Browse Definitions :
Definition

site scraper

Contributor(s): Ivy Wigmore

A site scraper is a type of software used to copy content from a website.

Site scrapers work similarly to web crawlers, which essentially perform the same function for the purposes of indexing websites. Web crawlers cover the whole Web, however, unlike site scrapers, which target user-specified websites.

Depending on the particular scraper program and user specifications, the software can download any data, including entire websites, and follow links to other content for further downloads. The data obtained may be saved as text, CSV, HTML or XML files; some scraper tools also enable export to a compatible database.

Content scraping has numerous legitimate purposes but is also often used for data theft and plagiarism. Websites featuring content scraped from other sites are called scraper sites.

Examples of site scrapers include Web Content Extractor, Wget, ScrapeGoat and Scraper, a Chrome extension.  

Asheesh Laroia explains web scraping in this video:

This was last updated in February 2014

Continue Reading About site scraper

Start the conversation

Send me notifications when other members comment.

Please create a username to comment.

SearchCompliance

  • risk assessment

    Risk assessment is the identification of hazards that could negatively impact an organization's ability to conduct business.

  • PCI DSS (Payment Card Industry Data Security Standard)

    The Payment Card Industry Data Security Standard (PCI DSS) is a widely accepted set of policies and procedures intended to ...

  • risk management

    Risk management is the process of identifying, assessing and controlling threats to an organization's capital and earnings.

SearchSecurity

SearchHealthIT

SearchDisasterRecovery

  • call tree

    A call tree is a layered hierarchical communication model that is used to notify specific individuals of an event and coordinate ...

  • Disaster Recovery as a Service (DRaaS)

    Disaster recovery as a service (DRaaS) is the replication and hosting of physical or virtual servers by a third party to provide ...

  • cloud disaster recovery (cloud DR)

    Cloud disaster recovery (cloud DR) is a combination of strategies and services intended to back up data, applications and other ...

SearchStorage

Close