Web scraping is the automated copying of content or data from websites so it can be used somewhere else. A program requests a page as a browser would, pulls out the pieces it wants - a price, a ranking, a headline - and stores them in a structured form. The Open Worldwide Application Security Project (OWASP), which lists scraping as automated threat OAT-011, describes it as collecting content and data "for use elsewhere", according to its definition. It exists because the web was built for people to read, not machines to query, and most sites offer no licensed feed for the data others want. Price comparison, search engine optimisation (SEO) tools and academic research depend on it, as does much of the text used to train artificial intelligence (AI) models.

How a scraper works

The basic cycle has four steps. It starts with a list of web addresses, or a seed list expanded by following links. The scraper then sends a Hypertext Transfer Protocol (HTTP) request for each page and receives HyperText Markup Language (HTML) in return. Parsing turns that HTML into a searchable tree, from which selectors pick specific elements such as a product's price. Finally, the extracted fields are cleaned and stored. The European Data Protection Board (EDPB) sets out the same sequence for AI training pipelines in its guidelines on scraping for generative AI.

Two distinctions drive the engineering. Static content sits in the HTML and needs one request. Dynamic content loads through JavaScript after the page opens, so the scraper must run a headless browser, a browser engine without a screen such as the open-source Puppeteer or Playwright. The EDPB also separates targeted scraping, bounded by fixed criteria such as one domain, from untargeted scraping by spiders that keep adding newly discovered links to their queue.

Scale brings resistance. Sites deploy rate limits, internet protocol (IP) address reputation lists, JavaScript challenges and CAPTCHA puzzles meant to separate people from programs. Scrapers respond by rotating requests across proxies. Data centre addresses are cheap and easy to spot; residential proxies route traffic through household connections instead. Google deployed SearchGuard, a JavaScript bot-detection layer, over its results pages in January 2025. Its lawsuit against SerpApi, filed on December 19, 2025, alleged that the Texas company worked around it and that its scraping volume rose by as much as 25,000% in two years.

The collecting side includes proxy vendors such as Bright Data, Oxylabs and Decodo; resellers such as SerpApi, which return search results through an application programming interface (API); SEO platforms; retailers; researchers; and AI developers. Opposite them stand publishers, marketplaces and social platforms, with security firms such as Cloudflare and DataDome filtering traffic on their behalf.

From web census to courtroom

Automated collection is nearly as old as the web. Matthew Gray's World Wide Web Wanderer, written at the Massachusetts Institute of Technology in June 1993, fetched pages simply to measure the web's size. In 2000 Judge Ronald Whyte granted eBay a preliminary injunction against Bidder's Edge, an auction aggregator sending 80,000 to 100,000 requests a day, holding that the load was likely a trespass to chattels, according to the Electronic Frontier Foundation's legal summary.

US law then narrowed. hiQ Labs, which analysed public LinkedIn profiles, persuaded the Ninth Circuit in 2019, and again in 2022 after the Supreme Court's 2021 Van Buren decision, that reading a publicly accessible site is not access "without authorization" under the Computer Fraud and Abuse Act (CFAA). Contract law reversed the result. In November 2022 the district court found LinkedIn's ban on scraping enforceable, and a consent judgment that December imposed $500,000 and a permanent injunction on hiQ, according to Morgan Lewis. Logged-out collection fared better: Meta's terms did not bar Bright Data's logged-off scraping of public pages, Judge Edward Chen ruled on January 23, 2024, according to Lowenstein Sandler, and X Corp's claims against the company were dismissed on May 10, 2024, according to Proskauer, Bright Data's counsel.

Europe fought on data protection grounds. Ireland's Data Protection Commission (DPC) fined Meta 265 million euros in November 2022 after a dataset built by exploiting Facebook's search and contact-import tools surfaced online, finding breaches of data protection by design under the General Data Protection Regulation (GDPR), according to the DPC. Twelve privacy authorities, including the UK's, stated on August 24, 2023 that publicly accessible personal data remains protected and that platforms carry duties against unlawful scraping, according to their joint statement. Copyright courts proved more permissive. Hamburg's Regional Court in September 2024, and the Higher Regional Court in December 2025, held that the non-profit LAION's image downloads for a training dataset fell within text and data mining exceptions, according to Hogan Lovells' Inside Tech Law.

Generative AI changed volume and motive. Elon Musk cited "extreme levels of data scraping" when Twitter restricted reading on July 1, 2023, according to Fortune, with initial caps of 6,000 posts a day for verified accounts and 600 for others, TechCrunch reported.

Why marketers depend on it

Much of the data marketers use is scraped. Rank trackers, keyword tools and share-of-voice dashboards read search engine results pages, directly or through vendors. When Google withdrew the num=100 parameter on September 14, 2025, a page of 100 results required ten requests, and Semrush described a tenfold rise in operating costs. Brands now scrape chatbots too: Decodo, a proxy vendor, placed ChatGPT third and Perplexity fifth in its 2026 ranking of the most-scraped sites, published on September 23 and drawn from its own customers' activity.

Publishers bear the cost. AI bots scraped Trusted Reviews 1.6 million times on August 16, 2025, a day that yielded 603 human visitors from AI platforms. Cloudflare Radar put bots at 57.4% of HTML requests in the week to June 5, 2026, with training crawlers alone at 50.6% of traffic.

The infrastructure also leaks into ad measurement. Spur found residential proxy code in more than 42% of LG webOS apps and over a quarter of Samsung Tizen apps, and Samsung banned such software development kits in August 2026. Verification vendors partly score traffic by checking whether an address belongs to a data centre. A household exit node passes that test.

Where the law and the defences fail

In the US the CFAA route has largely closed for public pages, leaving contract, trespass and unfair competition claims that must survive preemption by copyright law. Copyright itself has not filled the gap. Judge Yvonne Gonzalez Rogers dismissed Google's anti-circumvention claims against SerpApi on July 20, 2026 under the Digital Millennium Copyright Act (DMCA), finding that SearchGuard guards advertising revenue rather than copyright. Google amended its complaint on August 10, and a hearing on SerpApi's second motion to dismiss is set for October 13, 2026, according to the court docket.

Consistency is disputed as well. SerpApi's first motion to dismiss called Google "the largest scraper on the planet". Google's amended complaint answers that its own crawling happens in accordance with permissions and instructions that websites convey. SerpApi has also filed antitrust counterclaims over Reddit's $60 million Google-only crawler deal, invoking the hiQ and Bright Data rulings.

Europe treats scraping chiefly as a personal data question. France's Commission Nationale de l'Informatique et des Libertes (CNIL) has said mass scraping for AI training typically fails the GDPR's reasonable expectations test. Yet the Digital Services Act (DSA) requires the largest platforms to let vetted researchers collect public data, and the European Commission's 120 million euro fine on X on December 5, 2025 covered terms that barred researchers from scraping. Regulators police scraping and require it at once.

Technical controls leak. DataDome found full bot protection on just 2.4% of popular websites, with two in three tested sites stopping nothing. Robots.txt, the file that states crawling preferences, is voluntary, and TollBit counted 15% of AI page fetchers in Europe reaching pages publishers had disallowed. Common Crawl, whose archive feeds AI training, captured full paywalled articles because its crawler never runs the code that checks subscriptions. And the EDPB notes that personal data cannot easily be deleted from a trained model.

Not the same as

Web crawling. Following links to discover and fetch pages, usually to build an index. Scraping extracts specific data for reuse. The two overlap, and mixed-use crawlers collect pages for several purposes at once.

API access. A structured feed offered on the provider's terms. Reddit's licensees, for instance, connect to a Compliance API that notifies them when users delete posts. Scraped copies receive no such signal.

Account aggregation. Tools that sign into a service with a customer's own credentials, common in personal finance and often called screen scraping. OWASP files them under a separate threat, OAT-020.

Scraper sites. Pages that republish copied material to sell advertising. The term describes a use of scraped content, not the technique.

Recent developments

The EDPB adopted Guidelines 03/2026 on July 7, 2026, concluding that consent will most probably not work as a legal basis and that a missing robots.txt file is not consent; consultation closes on October 30. In California, Reddit's case against Anthropic returned to state court under a March 28, 2026 order finding its five contract and tort claims were not preempted by copyright. The hearing transcript records Anthropic's counsel, Ragesh Tangri, agreeing that Reddit's terms admit bots: "Where it draws the line is scraping." The order settles the forum, not the merits.

Infrastructure is moving too. Cloudflare, which opened pay per crawl with HTTP 402 Payment Required responses on July 1, 2025, shifted toward paying publishers when content appears in AI answers a year later. Since September 15, 2026, its network has blocked training and agent crawlers by default on ad-carrying pages of newly onboarded domains. Microsoft added an AI Scrape-to-Referral Ratio card to Clarity on August 13, 2026.

Timeline

  • June 1993: Matthew Gray's World Wide Web Wanderer, among the first web robots, begins measuring the size of the web
  • December 10, 1999: eBay sues auction aggregator Bidder's Edge over automated queries
  • May 24, 2000: Judge Ronald Whyte grants eBay a preliminary injunction on trespass to chattels grounds
  • 2004: Leonard Richardson releases Beautiful Soup, a Python library for parsing scraped HTML
  • 2014: Luminati begins selling access to Hola VPN users as exit nodes, the first large commercial residential proxy network
  • May 2017: LinkedIn sends hiQ Labs a cease-and-desist letter
  • September 2019: The Ninth Circuit holds that scraping public LinkedIn profiles likely does not breach the CFAA
  • June 3, 2021: The Supreme Court narrows the CFAA in Van Buren v. United States
  • April 2022: The Ninth Circuit reaffirms its hiQ ruling after remand
  • November 4, 2022: The district court finds hiQ breached LinkedIn's User Agreement
  • November 25, 2022: The Irish DPC fines Meta 265 million euros over the Facebook data scraping inquiry
  • December 8, 2022: A consent judgment imposes $500,000 and a permanent injunction on hiQ
  • July 1, 2023: Twitter caps daily post views, citing data scraping
  • August 24, 2023: Twelve data protection authorities issue a joint statement on data scraping
  • January 23, 2024: Judge Edward Chen rules Meta's terms do not bar Bright Data's logged-off scraping
  • May 10, 2024: X Corp's scraping claims against Bright Data are dismissed
  • September 27, 2024: Hamburg Regional Court rules LAION's dataset scraping falls within a text and data mining exception
  • October 28, 2024: The data protection authorities publish a concluding joint statement on scraping
  • January 2025: Google deploys SearchGuard to block automated access to search results
  • July 1, 2025: Cloudflare opens pay per crawl in private beta
  • August 6, 2025: Drop Site News publishes a leaked Meta list of sites targeted for scraping
  • August 16, 2025: AI bots scrape Trusted Reviews 1.6 million times in a single day
  • September 14, 2025: Google withdraws the num=100 search parameter
  • October 22, 2025: Reddit sues SerpApi, Oxylabs, AWMProxy and Perplexity in New York
  • December 5, 2025: The European Commission fines X 120 million euros under the DSA, including for barring researcher scraping
  • December 10, 2025: The Hanseatic Higher Regional Court dismisses the appeal against LAION
  • December 19, 2025: Google sues SerpApi under the DMCA over SearchGuard circumvention
  • March 28, 2026: A federal judge remands Reddit's five scraping claims against Anthropic to state court
  • June 5, 2026: Cloudflare Radar records bots at 57.4% of HTML requests over seven days
  • July 1, 2026: Cloudflare shifts from per-crawl charging toward per-answer payments
  • July 7, 2026: The EDPB adopts Guidelines 03/2026 on web scraping for generative AI
  • July 20, 2026: Google's DMCA claims against SerpApi are dismissed
  • August 3, 2026: Samsung restricts residential proxy SDKs on Tizen
  • August 10, 2026: Google files an amended complaint against SerpApi
  • August 13, 2026: Microsoft Clarity adds an AI Scrape-to-Referral Ratio card
  • August 28, 2026: SerpApi files antitrust counterclaims against Reddit
  • September 15, 2026: Cloudflare begins default blocking of training and agent crawlers on ad-carrying pages of new domains
  • September 23, 2026: Decodo publishes its third annual most-scraped websites ranking
  • October 13, 2026: Scheduled hearing on SerpApi's motion to dismiss Google's amended complaint
  • October 30, 2026: The EDPB's public consultation on Guidelines 03/2026 is due to close

Summary

Who: Proxy and data vendors, search data resellers, SEO and competitive intelligence platforms, retailers, researchers and AI developers do the scraping. Publishers, marketplaces and social platforms defend against it, with Cloudflare, DataDome and similar firms filtering traffic. Courts, data protection authorities and the European Commission set the limits.

What: The automated extraction of content or data from web pages for reuse elsewhere, using HTTP requests, HTML parsing and, for pages built with JavaScript, headless browsers, often routed through proxy networks to avoid detection.

When: Web robots date from June 1993. The first major lawsuit reached an injunction in May 2000, US law on public pages took shape between 2019 and 2024, and AI training pushed scraping into regulators' guidance and new litigation from 2023, with EDPB guidelines under consultation until October 30, 2026.

Where: Across the open web, with the heaviest activity on search results, social video, marketplaces and, since 2026, AI chatbots. Key disputes run in federal courts in California and New York, California state court, and before EU regulators.

Why: Most websites offer no licensed feed for the data others want, so scraping fills the gap for price monitoring, search measurement and model training. It also shifts costs onto the sites being read, blurs the line between human and machine traffic in advertising measurement, and leaves ownership of public data to be settled case by case.