What is site scraper? - Definition from WhatIs.com
Part of the Content management glossary:

A site scraper is a type of software used to copy content from a website.

Site scrapers work similarly to web crawlers, which essentially perform the same function for the purposes of indexing websites. Web crawlers cover the whole Web, however, unlike site scrapers, which target user-specified websites.

Depending on the particular scraper program and user specifications, the software can download any data, including entire websites, and follow links to other content for further downloads. The data obtained may be saved as text, CSV, HTML or XML files; some scraper tools also enable export to a compatible database.

Content scraping has numerous legitimate purposes but is also often used for data theft and plagiarism. Websites featuring content scraped from other sites are called scraper sites.

Examples of site scrapers include Web Content Extractor, Wget, ScrapeGoat and Scraper, a Chrome extension.  

Asheesh Laroia explains web scraping in this video:

This was last updated in February 2014
Contributor(s): Ivy Wigmore
Posted by: Margaret Rouse

Related Terms

Definitions

  • enterprise search

    - Enterprise search is the organized retrieval of structured and unstructured data within an organization. (WhatIs.com)

  • semi-structured data

    - Semi-structured data is data that has not been organized into a specialized format, such as a table, a record, an array or a tree but that nevertheless has associated information, such as metadata,... (WhatIs.com)

  • recommendation engine

    - Recommendation engines are common among online retail websites, such as Amazon. Also known as recommender systems, these applications suggest products (or something else a visitor might search for,... (WhatIs.com)

Glossaries

  • Content management

    - Terms related to content management, including definitions about enterprise content management and words and phrases about content management applications (CMA) and content management systems (CMS).

  • Internet applications

    - This WhatIs.com glossary contains terms related to Internet applications, including definitions about Software as a Service (SaaS) delivery models and words and phrases about web sites, e-commerce ...

Ask a Question About site scraperPowered by ITKnowledgeExchange.com

Get answers from your peers on your most technical challenges

Tech TalkComment

Share
Comments

    Results

    Contribute to the conversation

    All fields are required. Comments will appear at the bottom of the article.