Browse Definitions :
Definition

Googlebot

Contributor(s): Matthew Haughn

Googlebot is a web crawling software search bot (also known as a spider or webcrawler) that gathers the web page information used to supply Google search engine results pages (SERP).

Googlebot collects documents from the web to build Google’s search index. Through constantly gathering documents, the software discovers new pages and updates to existing pages. Googlebot uses a distributed design spanning many computers so it can grow as the web does.

The webcrawler uses algorithms to determine what sites to browse, what rates to browse at and how many pages to fetch from. Googlebot begins with a list generated from previous sessions. This list is then augmented by the sitemaps provided by webmasters. The software crawls all linked elements in the webpages it browses, noting new sites, updates to sites and dead links. The information gathered is used to update Google’s index of the web.

Googlebot creates an index within the limitations set forth by webmasters in their robots.txt files. Should a webmaster wish to keep pages hidden from Google search, for example, he can block Googlebot in a robots.txt file at the top-level folder of the site. To prevent Googlebot from following any links on a given page of a site, he can include the nofollow meta tag; to prevent the bot from following individual links, the webmaster can add rel="nofollow" to the links themselves.

A site’s webmaster might detect visits every few seconds from computers at google.com, showing the user-agent Googlebot. Generally, Google tries to index as much of a site as it can without overwhelming the site’s bandwidth. If a webmaster finds that Googlebot is using too much bandwidth, they can set a rate on Google’s search console homepage that will remain in effect for 90 days.

Presenting at the 2011 SearchLove conference, Josh Giardino claimed that Googlebot is actually the Chrome browser. That would mean that Googlebot has not only the ability to browse pages in text, as crawlers do, but can also run scripts and media as web browsers do. That capacity could allow Googlebot to find hidden information and perform other tasks that are not acknowledged by Google. Giardino went so far as to say that Googlebot may be the original reason that the company created Chrome.

This was last updated in June 2017

Continue Reading About Googlebot

Start the conversation

Send me notifications when other members comment.

Please create a username to comment.

-ADS BY GOOGLE

File Extensions and File Formats

Powered by:

SearchCompliance

  • California Consumer Privacy Act (CCPA)

    The California Consumer Privacy Act (CCPA) is legislation in the state of California that supports an individual's right to ...

  • compliance audit

    A compliance audit is a comprehensive review of an organization's adherence to regulatory guidelines.

  • regulatory compliance

    Regulatory compliance is an organization's adherence to laws, regulations, guidelines and specifications relevant to its business...

SearchSecurity

  • endpoint detection and response (EDR)

    Endpoint detection and response (EDR) is a category of tools and technology used for protecting computer hardware devices–called ...

  • ransomware

    Ransomware is a subset of malware in which the data on a victim's computer is locked, typically by encryption, and payment is ...

  • single sign-on (SSO)

    Single sign-on (SSO) is a session and user authentication service that permits an end user to enter one set of login credentials ...

SearchHealthIT

SearchDisasterRecovery

  • disaster recovery team

    A disaster recovery team is a group of individuals focused on planning, implementing, maintaining, auditing and testing an ...

  • cloud insurance

    Cloud insurance is any type of financial or data protection obtained by a cloud service provider. 

  • business continuity software

    Business continuity software is an application or suite designed to make business continuity planning/business continuity ...

SearchStorage

  • blockchain storage

    Blockchain storage is a way of saving data in a decentralized network which utilizes the unused hard disk space of users across ...

  • disk mirroring (RAID 1)

    RAID 1 is one of the most common RAID levels and the most reliable. Data is written to two places simultaneously, so if one disk ...

  • RAID controller

    A RAID controller is a hardware device or software program used to manage hard disk drives (HDDs) or solid-state drives (SSDs) in...

Close