Regurgitation is the reproduction of training data, word for word or nearly so, in the output of a generative model. Nothing is fetched from a database at request time. The passage comes back out of the model's parameters, reconstructed token by token from what the system absorbed months or years earlier. The term matters because the behaviour contradicts the standard account of how these systems work. Developers have argued that training extracts statistical patterns rather than storing copies, and Andreessen Horowitz told the US Copyright Office in October 2023 that research showed extremely small rates of memorisation. Regurgitation is the observable counter-example.

The word reached advertising and publishing through litigation rather than research. It now appears in court filings as a quantified term with a measured rate, and that number is being used to decide whether the content underneath the generative economy has to be paid for.

How memorisation becomes output

A language model predicts the next token given everything before it. Where a sequence appeared in training data often enough, or in a sufficiently distinctive form, the parameters encode its continuation so sharply that the highest-probability token at every position reconstructs the original. No architectural boundary separates a learned generalisation from a learned passage. The US Copyright Office, in its May 2025 report, cited researchers A. Feder Cooper and James Grimmelmann for the point that learned patterns run from highly abstract to highly specific, and that where the pattern is specific enough, the pattern is the memorised training data.

Three variables drive the rate. Duplication is the strongest: a sentence syndicated across thousands of scraped pages is far more likely to come back than one published once. Model size is the second, and Nicholas Carlini and co-authors showed in 2022 that the six-billion-parameter GPT-J model had memorised at least 1% of its training corpus. Context length is the third, since the more of a passage supplied as a prefix, the more reliably the rest follows.

Researchers separate discoverable memorisation, which asks whether a known passage can be recovered when its opening is supplied, from extractable memorisation, which asks what an adversary can pull out without knowing the training set. Litigation runs on the first. Security research runs on the second.

Two mechanisms, one word

Verbatim output has a second source that looks identical. In retrieval-augmented systems a document is fetched at query time and inserted into the context window, so nothing needs to be memorised. Cable News Network's complaint against Perplexity, filed on May 28, 2026 over more than 17,000 works, describes the Comet browser assistant reproducing text from a paywalled article while the subscription prompt sat on the same screen.

The distinction is legally significant and commercially invisible. Weight-based regurgitation implicates the training corpus and the model. Retrieval-based reproduction implicates the crawler, the index and the product. In the consolidated New York Times proceedings the publishers press copying claims at five stages, treating acquisition, training, grounding and outputs as distinct acts.

Extraction techniques and what they recovered

The first practical demonstration came from Carlini and eleven co-authors at USENIX Security in 2021. Working against GPT-2, they generated 1,800 candidate sequences and confirmed more than 600 as verbatim training samples, including personal contact details.

Aligned chatbots resisted those methods, so Milad Nasr, Carlini and colleagues published a divergence attack on November 28, 2023. Instructing ChatGPT to repeat a single word indefinitely made it abandon its dialogue behaviour and fall back toward raw language modelling, emitting training data at 150 times the normal rate. A draft Council of Europe expert report, now heading for adoption in November 2026, cites a study finding that certain trigger words raised emission of memorised data up to 164 times above a neutral control, and that extracting email signatures cost around $200.

Scale changed in January 2026, when Stanford researchers extracted near-complete books from four production systems. Claude 3.7 Sonnet reproduced 95.8% of Harry Potter and the Sorcerer's Stone after 258 jailbreak attempts, and 97.5% of The Great Gatsby, at $55 to $135 per book. Gemini 2.5 Pro returned 76.8% of the same title with no jailbreaking at all, for $2.44, and Grok 3 managed 70.3% for $8.16. GPT-4.1 required 5,179 attempts and still stopped near 4%, refusing at chapter boundaries. The longest contiguous block ran to 9,070 words, against the roughly 38-word sequences typical of academic studies.

Counting it in court

The number became the centre of a case on September 4, 2026, when OpenAI, Microsoft and five news organisations each moved for summary judgment before Judge Sidney H. Stein. Searching the 20 million conversation logs produced in discovery, OpenAI's expert reported 24 instances of verbatim regurgitation, a rate of 0.00012%. The plaintiffs' own experts landed in the same range, from 0.00011% for the Times to 0.000002% for The Intercept, and the longest passages ran to 29 and 43 words. Prompts covering 710 asserted works produced more than 68,000 outputs with none at all.

Adversarial testing produced larger figures. The plaintiffs' experts used more than 250 million prompts against non-public models and API-only configurations, triggering partial regurgitation of fewer than 2.6% of asserted works and extracting an average below 4.4% of each. OpenAI set that against the 16% recovered in the Google Books case, where the Second Circuit found fair use.

Thresholds are contested too. The Times and the Daily News plaintiffs seek judgment only where 80% of three-word sequences match and the texts share a sixteen-word run. Everything below that line goes to a jury.

Why it reaches the marketing industry

Three consequences follow for anyone buying or selling media. The first is the price of content. If regurgitation is rare enough to be dismissed, the fourth fair use factor weakens and the licensing market publishers are trying to build loses its floor. The New York Times spends close to $2 billion a year producing about half a million works, a figure disclosed in August 2026 while the company sued three artificial intelligence companies and licensed to a fourth.

The second is liability for generated creative. In November 2025 the Munich Regional Court I found that OpenAI's models contained reproducible lyrics from nine German songs, holding that memorisation amounts to a fixation of the work in the parameters. According to the European Commission's intellectual property helpdesk, the court rejected the argument that users prompting the system are responsible for its output. The same chamber ruled against Suno on July 31, 2026 and set a penalty of 250,000 euros per breach, noting that the defendant had been aware of memorisation since the 2021 Carlini paper. Liability sits with the model provider, which reaches every agency shipping generative output into paid media.

The third is inventory. The publishers' filings describe synthetic sites recycling professional journalism at roughly $6,800 per million articles, competing for the same programmatic budgets as newsroom material. Integral Ad Science flagged machine-generated sites as an ad quality threat in July 2025.

Limits, criticisms and open disputes

Every reported rate is a function of how hard someone tried. Low organic rates and high adversarial rates are both true, and each side cites the measurement that suits it. Whether hostile prompting describes a real harm or a manufactured one remains unresolved.

Mitigations leak. The Council of Europe report records that deduplicating training data reduces memorisation but introduces security side channels, that shrinking the context window constrains the retrieval applications enterprises actually deploy, and that filtering memorised sequences out of output is ineffective.

Courts disagree on the underlying technical question. The High Court of England and Wales held in November 2025 that Stable Diffusion does not store copies of training images in its weights, granting Getty Images only a narrow trademark finding on watermarks. Munich reached the opposite conclusion a week later. Neither ruling has an appellate answer.

A privacy dimension runs alongside the copyright one. Research published in July 2025 concluded that large language models qualify as personal data under European Union rules when they memorise training information, with rates estimated between 0.1% and 10%. France's CNIL requires developers to query their own models against selected request lists to establish what personal data has been retained.

Auditing any of it depends on disclosure that does not exist. A House of Lords committee demanded in March 2026 that AI firms reveal training data, and the TRAIN Act, introduced in Congress on January 22, 2026, would give copyright owners subpoena rights over training records.

Disambiguation

Memorisation is the internal condition, the retention of specific training content in model parameters. Regurgitation is the observable behaviour proving it occurred.

Hallucination is the opposite failure. A hallucinating model produces fluent text with no source at all, while a regurgitating model produces text whose source is too exact.

Regurgitative training means training a model on output generated by other models, a separate problem associated with degraded quality across generations.

Grounded reproduction covers verbatim text arriving through retrieval, crawling or a browsing agent rather than from weights. Nothing was memorised; a document was copied into the context window at request time.

Recent developments

The measurement fight has become the case. Both sides in the New York Times litigation now argue from regurgitation statistics rather than over whether reproduction happens at all, and the Department of Justice filed a statement supporting OpenAI's fair use defence on September 3, 2026, a day before the motions landed. Europe runs the other way, with two Munich judgments treating memorisation as reproduction and the Council of Europe preparing guidelines that treat it as a data protection failure. As of September 2026 no appellate court has ruled on whether a model able to reproduce a work contains it.

Timeline

  • 2019: Carlini and co-authors publish The Secret Sharer, establishing unintended memorisation in neural networks
  • December 2020: The GPT-2 extraction paper is posted, later presented at USENIX Security in August 2021
  • 2022: Quantifying Memorization Across Neural Language Models shows GPT-J memorised at least 1% of The Pile
  • November 28, 2023: The divergence attack paper demonstrates training data emission from ChatGPT at 150 times the normal rate
  • December 2023: The New York Times sues OpenAI and Microsoft, attaching over 100 examples of verbatim reproduction
  • May 2025: The US Copyright Office publishes its report on generative AI training, addressing memorisation directly
  • June 2025: A federal court finds Meta's use of books for training transformative, citing effective output filters
  • July 2025: Research concludes that memorising models qualify as personal data under EU law
  • September 2025: Anthropic agrees to a $1.5 billion settlement in an authors' copyright case
  • November 2025: The High Court of England and Wales holds that Stable Diffusion does not store copies of training images
  • November 11, 2025: The Munich Regional Court I finds that OpenAI's models reproduce nine German song lyrics
  • November 13, 2025: A US court orders production of 20 million ChatGPT conversation logs
  • January 6, 2026: Stanford researchers publish near-complete book extractions from four production models
  • January 22, 2026: The TRAIN Act is introduced in the US House of Representatives
  • May 28, 2026: CNN sues Perplexity over more than 17,000 works, including verbatim paywalled text
  • July 31, 2026: The Munich Regional Court I rules against Suno, setting a 250,000 euro penalty per breach
  • September 4, 2026: OpenAI, Microsoft and five news organisations file cross-motions for summary judgment

Summary

Who. Model developers including OpenAI, Anthropic, Google, xAI and Perplexity operate the systems that produce regurgitated output. Publishers, record labels, collecting societies and authors are the parties measuring it. Security researchers at Google, Stanford, ETH Zurich and elsewhere designed the extraction methods, and courts in New York, Munich and London are deciding what the results mean.

What. Regurgitation is verbatim or near-verbatim reproduction of training data in a model's output. It is the observable evidence that memorisation occurred during training, and it is distinct from retrieval-based reproduction, where the source document passes through the context window at query time.

When. The behaviour was documented academically from 2019 onward and demonstrated at scale against production systems in November 2023. It became a litigated quantity in December 2023 and a decided one in Munich in November 2025. Cross-motions turning on the measured rate were filed in New York on September 4, 2026.

Where. Regurgitation surfaces in chat interfaces, model APIs, browsing agents and any advertising or content tool built on a general-purpose model. It is measured in conversation logs produced in discovery, in adversarial extraction experiments and in prompt-and-output comparisons run by courts.

Why. The rate determines whether copying that happens rarely is treated as an edge case or as proof that models contain the works they were trained on. That answer sets the price of content for every generative system built on the open web, decides who carries liability when generated creative reproduces protected material, and shapes the economics of the inventory competing for programmatic spend.