Web scraping is the automated copying of content or data from websites so it can be used somewhere else. A program requests a page as a browser would, pulls out the pieces it wants - a price, a ranking, a headline - and stores them in a structured form. The Open Worldwide Application Security Project (OWASP), which lists scraping as automated threat OAT-011, describes it as collecting content and data "for use elsewhere", according to its definition. It exists because the web was built for people to read, not machines to query, and most sites offer no licensed feed for the data others want. Price comparison, search engine optimisation (SEO) tools and academic research depend on it, as does much of the text used to train artificial intelligence (AI) models.
How a scraper works
The basic cycle has four steps. It starts with a list of web addresses, or a seed list expanded by following links. The scraper then sends a Hypertext Transfer Protocol (HTTP) request for each page and receives HyperText Markup Language (HTML) in return. Parsing turns that HTML into a searchable tree, from which selectors pick specific elements such as a product's price. Finally, the extracted fields are cleaned and stored. The European Data Protection Board (EDPB) sets out the same sequence for AI training pipelines in its guidelines on scraping for generative AI.
Two distinctions drive the engineering. Static content sits in the HTML and needs one request. Dynamic content loads through JavaScript after the page opens, so the scraper must run a headless browser, a browser engine without a screen such as the open-source Puppeteer or Playwright. The EDPB also separates targeted scraping, bounded by fixed criteria such as one domain, from untargeted scraping by spiders that keep adding newly discovered links to their queue.
Scale brings resistance. Sites deploy rate limits, internet protocol (IP) address reputation lists, JavaScript challenges and CAPTCHA puzzles meant to separate people from programs. Scrapers respond by rotating requests across proxies. Data centre addresses are cheap and easy to spot; residential proxies route traffic through household connections instead. Google deployed SearchGuard, a JavaScript bot-detection layer, over its results pages in January 2025. Its lawsuit against SerpApi, filed on December 19, 2025, alleged that the Texas company worked around it and that its scraping volume rose by as much as 25,000% in two years.
The collecting side includes proxy vendors such as Bright Data, Oxylabs and Decodo; resellers such as SerpApi, which return search results through an application programming interface (API); SEO platforms; retailers; researchers; and AI developers. Opposite them stand publishers, marketplaces and social platforms, with security firms such as Cloudflare and DataDome filtering traffic on their behalf.
From web census to courtroom
Automated collection is nearly as old as the web. Matthew Gray's World Wide Web Wanderer, written at the Massachusetts Institute of Technology in June 1993, fetched pages simply to measure the web's size. In 2000 Judge Ronald Whyte granted eBay a preliminary injunction against Bidder's Edge, an auction aggregator sending 80,000 to 100,000 requests a day, holding that the load was likely a trespass to chattels, according to the Electronic Frontier Foundation's legal summary.
US law then narrowed. hiQ Labs, which analysed public LinkedIn profiles, persuaded the Ninth Circuit in 2019, and again in 2022 after the Supreme Court's 2021 Van Buren decision, that reading a publicly accessible site is not access "without authorization" under the Computer Fraud and Abuse Act (CFAA). Contract law reversed the result. In November 2022 the district court found LinkedIn's ban on scraping enforceable, and a consent judgment that December imposed $500,000 and a permanent injunction on hiQ, according to Morgan Lewis. Logged-out collection fared better: Meta's terms did not bar Bright Data's logged-off scraping of public pages, Judge Edward Chen ruled on January 23, 2024, according to Lowenstein Sandler, and X Corp's claims against the company were dismissed on May 10, 2024, according to Proskauer, Bright Data's counsel.
Europe fought on data protection grounds. Ireland's Data Protection Commission (DPC) fined Meta 265 million euros in November 2022 after a dataset built by exploiting Facebook's search and contact-import tools surfaced online, finding breaches of data protection by design under the General Data Protection Regulation (GDPR), according to the DPC. Twelve privacy authorities, including the UK's, stated on August 24, 2023 that publicly accessible personal data remains protected and that platforms carry duties against unlawful scraping, according to their joint statement. Copyright courts proved more permissive. Hamburg's Regional Court in September 2024, and the Higher Regional Court in December 2025, held that the non-profit LAION's image downloads for a training dataset fell within text and data mining exceptions, according to Hogan Lovells' Inside Tech Law.
Generative AI changed volume and motive. Elon Musk cited "extreme levels of data scraping" when Twitter restricted reading on July 1, 2023, according to Fortune, with initial caps of 6,000 posts a day for verified accounts and 600 for others, TechCrunch reported.
Why marketers depend on it
Much of the data marketers use is scraped. Rank trackers, keyword tools and share-of-voice dashboards read search engine results pages, directly or through vendors. When Google withdrew the num=100 parameter on September 14, 2025, a page of 100 results required ten requests, and Semrush described a tenfold rise in operating costs. Brands now scrape chatbots too: Decodo, a proxy vendor, placed ChatGPT third and Perplexity fifth in its 2026 ranking of the most-scraped sites, published on September 23 and drawn from its own customers' activity.
Publishers bear the cost. AI bots scraped Trusted Reviews 1.6 million times on August 16, 2025, a day that yielded 603 human visitors from AI platforms. Cloudflare Radar put bots at 57.4% of HTML requests in the week to June 5, 2026, with training crawlers alone at 50.6% of traffic.
The infrastructure also leaks into ad measurement. Spur found residential proxy code in more than 42% of LG webOS apps and over a quarter of Samsung Tizen apps, and Samsung banned such software development kits in August 2026. Verification vendors partly score traffic by checking whether an address belongs to a data centre. A household exit node passes that test.
Where the law and the defences fail
In the US the CFAA route has largely closed for public pages, leaving contract, trespass and unfair competition claims that must survive preemption by copyright law. Copyright itself has not filled the gap. Judge Yvonne Gonzalez Rogers dismissed Google's anti-circumvention claims against SerpApi on July 20, 2026 under the Digital Millennium Copyright Act (DMCA), finding that SearchGuard guards advertising revenue rather than copyright. Google amended its complaint on August 10, and a hearing on SerpApi's second motion to dismiss is set for October 13, 2026, according to the court docket.
Consistency is disputed as well. SerpApi's first motion to dismiss called Google "the largest scraper on the planet". Google's amended complaint answers that its own crawling happens in accordance with permissions and instructions that websites convey. SerpApi has also filed antitrust counterclaims over Reddit's $60 million Google-only crawler deal, invoking the hiQ and Bright Data rulings.
Europe treats scraping chiefly as a personal data question. France's Commission Nationale de l'Informatique et des Libertes (CNIL) has said mass scraping for AI training typically fails the GDPR's reasonable expectations test. Yet the Digital Services Act (DSA) requires the largest platforms to let vetted researchers collect public data, and the European Commission's 120 million euro fine on X on December 5, 2025 covered terms that barred researchers from scraping. Regulators police scraping and require it at once.
Technical controls leak. DataDome found full bot protection on just 2.4% of popular websites, with two in three tested sites stopping nothing. Robots.txt, the file that states crawling preferences, is voluntary, and TollBit counted 15% of AI page fetchers in Europe reaching pages publishers had disallowed. Common Crawl, whose archive feeds AI training, captured full paywalled articles because its crawler never runs the code that checks subscriptions. And the EDPB notes that personal data cannot easily be deleted from a trained model.
Not the same as
Web crawling. Following links to discover and fetch pages, usually to build an index. Scraping extracts specific data for reuse. The two overlap, and mixed-use crawlers collect pages for several purposes at once.
API access. A structured feed offered on the provider's terms. Reddit's licensees, for instance, connect to a Compliance API that notifies them when users delete posts. Scraped copies receive no such signal.
Account aggregation. Tools that sign into a service with a customer's own credentials, common in personal finance and often called screen scraping. OWASP files them under a separate threat, OAT-020.
Scraper sites. Pages that republish copied material to sell advertising. The term describes a use of scraped content, not the technique.
Recent developments
The EDPB adopted Guidelines 03/2026 on July 7, 2026, concluding that consent will most probably not work as a legal basis and that a missing robots.txt file is not consent; consultation closes on October 30. In California, Reddit's case against Anthropic returned to state court under a March 28, 2026 order finding its five contract and tort claims were not preempted by copyright. The hearing transcript records Anthropic's counsel, Ragesh Tangri, agreeing that Reddit's terms admit bots: "Where it draws the line is scraping." The order settles the forum, not the merits.
Infrastructure is moving too. Cloudflare, which opened pay per crawl with HTTP 402 Payment Required responses on July 1, 2025, shifted toward paying publishers when content appears in AI answers a year later. Since September 15, 2026, its network has blocked training and agent crawlers by default on ad-carrying pages of newly onboarded domains. Microsoft added an AI Scrape-to-Referral Ratio card to Clarity on August 13, 2026.
Timeline
- June 1993: Matthew Gray's World Wide Web Wanderer, among the first web robots, begins measuring the size of the web
- December 10, 1999: eBay sues auction aggregator Bidder's Edge over automated queries
- May 24, 2000: Judge Ronald Whyte grants eBay a preliminary injunction on trespass to chattels grounds
- 2004: Leonard Richardson releases Beautiful Soup, a Python library for parsing scraped HTML
- 2014: Luminati begins selling access to Hola VPN users as exit nodes, the first large commercial residential proxy network
- May 2017: LinkedIn sends hiQ Labs a cease-and-desist letter
- September 2019: The Ninth Circuit holds that scraping public LinkedIn profiles likely does not breach the CFAA
- June 3, 2021: The Supreme Court narrows the CFAA in Van Buren v. United States
- April 2022: The Ninth Circuit reaffirms its hiQ ruling after remand
- November 4, 2022: The district court finds hiQ breached LinkedIn's User Agreement
- November 25, 2022: The Irish DPC fines Meta 265 million euros over the Facebook data scraping inquiry
- December 8, 2022: A consent judgment imposes $500,000 and a permanent injunction on hiQ
- July 1, 2023: Twitter caps daily post views, citing data scraping
- August 24, 2023: Twelve data protection authorities issue a joint statement on data scraping
- January 23, 2024: Judge Edward Chen rules Meta's terms do not bar Bright Data's logged-off scraping
- May 10, 2024: X Corp's scraping claims against Bright Data are dismissed
- September 27, 2024: Hamburg Regional Court rules LAION's dataset scraping falls within a text and data mining exception
- October 28, 2024: The data protection authorities publish a concluding joint statement on scraping
- January 2025: Google deploys SearchGuard to block automated access to search results
- July 1, 2025: Cloudflare opens pay per crawl in private beta
- August 6, 2025: Drop Site News publishes a leaked Meta list of sites targeted for scraping
- August 16, 2025: AI bots scrape Trusted Reviews 1.6 million times in a single day
- September 14, 2025: Google withdraws the num=100 search parameter
- October 22, 2025: Reddit sues SerpApi, Oxylabs, AWMProxy and Perplexity in New York
- December 5, 2025: The European Commission fines X 120 million euros under the DSA, including for barring researcher scraping
- December 10, 2025: The Hanseatic Higher Regional Court dismisses the appeal against LAION
- December 19, 2025: Google sues SerpApi under the DMCA over SearchGuard circumvention
- March 28, 2026: A federal judge remands Reddit's five scraping claims against Anthropic to state court
- June 5, 2026: Cloudflare Radar records bots at 57.4% of HTML requests over seven days
- July 1, 2026: Cloudflare shifts from per-crawl charging toward per-answer payments
- July 7, 2026: The EDPB adopts Guidelines 03/2026 on web scraping for generative AI
- July 20, 2026: Google's DMCA claims against SerpApi are dismissed
- August 3, 2026: Samsung restricts residential proxy SDKs on Tizen
- August 10, 2026: Google files an amended complaint against SerpApi
- August 13, 2026: Microsoft Clarity adds an AI Scrape-to-Referral Ratio card
- August 28, 2026: SerpApi files antitrust counterclaims against Reddit
- September 15, 2026: Cloudflare begins default blocking of training and agent crawlers on ad-carrying pages of new domains
- September 23, 2026: Decodo publishes its third annual most-scraped websites ranking
- October 13, 2026: Scheduled hearing on SerpApi's motion to dismiss Google's amended complaint
- October 30, 2026: The EDPB's public consultation on Guidelines 03/2026 is due to close
Related PPC Land coverage
- EDPB blocks AI firms from using consent as an excuse to scrape - The July 2026 guidelines, their four-step scraping sequence and the legitimate interest test.
- Explaining residential proxy - How household exit nodes carry scraping traffic and why they undermine IP-based filtering.
- Google sues SerpApi over search scraping in copyright lawsuit - The December 2025 complaint, SearchGuard and the claimed 25,000% rise in scraping volume.
- SEO industry adapts as Google forces 10x cost increase on tracking platforms - How the num=100 withdrawal multiplied the cost of collecting search data.
- Seven of ten spots turn over as ChatGPT enters Decodo's 2026 scraping ranking - The September 2026 ranking showing AI answer engines among the most-scraped sites.
- UK publishers bill AI scrapers 500 pounds per article using county courts - The Trusted Reviews scraping surge and the contract route publishers are testing.
- Bots now outnumber humans on the web - and most aren't here to search - Cloudflare Radar data putting bots at 57.4% of HTML requests.
- Samsung bans proxy SDKs as a quarter of Tizen apps route strangers' traffic - Residential proxy code in smart TV apps and its effect on traffic verification.
- Google loses DMCA bid to treat search scraping like DVD piracy - The July 2026 dismissal finding SearchGuard protects ad revenue, not copyright.
- SerpApi files motion to dismiss Google's DMCA scraping lawsuit - SerpApi's February 2026 arguments on standing, circumvention and Google's own crawling.
- SerpApi faces revived Google scraping claims built on Reddit licensing terms - Google's August 2026 amended complaint and its framing of permission-based crawling.
- Reddit faces antitrust counterclaims over $60m Google-only crawler deal - SerpApi's counterclaims invoking hiQ and X Corp v. Bright Data.
- GDPR's AI training legal battle: regulators converge but still clash - How 19 regulatory guidelines treat scraping, including the CNIL's reasonable expectations position.
- European Commission fines X 120 million euros for transparency violations - The first DSA non-compliance decision, covering terms that blocked researcher scraping.
- Full bot protection drops to 2.4% of popular websites, DataDome finds - Benchmark results showing how few sites stop automated requests.
- ChatGPT ads reach Europe as its own crawler ignores publisher blocks - TollBit's finding that 15% of AI page fetchers reached disallowed URLs, and Cloudflare's September 15 default.
- Common Crawl supplies paywalled content to AI companies despite publisher objections - How a non-profit archive captured paywalled articles for AI training.
- Explaining mixed-use crawler - Crawlers that collect pages for search, AI training and other purposes at once.
- Anthropic loses bid to keep Reddit's 5 scraping claims in federal court - The remand order, the Compliance API and the hearing exchange on where scraping begins.
- Cloudflare ties AI payouts to citations as 50% of crawls waste - The July 2026 move from Pay Per Crawl toward Pay Per Use.
- Microsoft Clarity card ranks which AI operators scrape most and refer least - The August 2026 dashboard card comparing scraping volume with referral traffic.
Summary
Who: Proxy and data vendors, search data resellers, SEO and competitive intelligence platforms, retailers, researchers and AI developers do the scraping. Publishers, marketplaces and social platforms defend against it, with Cloudflare, DataDome and similar firms filtering traffic. Courts, data protection authorities and the European Commission set the limits.
What: The automated extraction of content or data from web pages for reuse elsewhere, using HTTP requests, HTML parsing and, for pages built with JavaScript, headless browsers, often routed through proxy networks to avoid detection.
When: Web robots date from June 1993. The first major lawsuit reached an injunction in May 2000, US law on public pages took shape between 2019 and 2024, and AI training pushed scraping into regulators' guidance and new litigation from 2023, with EDPB guidelines under consultation until October 30, 2026.
Where: Across the open web, with the heaviest activity on search results, social video, marketplaces and, since 2026, AI chatbots. Key disputes run in federal courts in California and New York, California state court, and before EU regulators.
Why: Most websites offer no licensed feed for the data others want, so scraping fills the gap for price monitoring, search measurement and model training. It also shifts costs onto the sites being read, blurs the line between human and machine traffic in advertising measurement, and leaves ownership of public data to be settled case by case.
Discussion