A Microsoft scientist wrote that people would soon regard large models "hoovering up" their work as "an astonishing theft of unprecedented proportions." That line now opens the 92-page brief five news publishers filed against Microsoft and OpenAI on September 17, 2026, alongside click-through data, copy counts and internal messages that were largely hidden when the motions were first lodged.

The New York Times Company, eight Daily News titles, Ziff Davis, the Center for Investigative Reporting and The Intercept filed the public version of their combined summary judgment brief in the Southern District of New York on September 17, 2026. The document, docketed as 1977-1 in the multidistrict litigation numbered 25-md-3143 before Judge Sidney H. Stein, asks the court to rule that OpenAI and Microsoft infringed copyright at five stages of building and running their artificial intelligence products, and that fair use excuses none of them.

The motions themselves are not new. All three parties moved for summary judgment on September 4, and PPC Land reported at the time that the public versions arrived with black bars over traffic figures, internal executive commentary and the fifth category of alleged copying. The September 17 filing still carries redactions, but far fewer. What has emerged is the publishers' account of what the two companies said to each other, in writing, about the content they were taking.

In Short

Five news organisations say OpenAI and Microsoft copied millions of their articles without paying, and a new court filing shows internal messages from both companies describing that copying in stark terms, including one calling it theft. This matters to anyone who publishes content online, because the case will help decide whether AI companies must pay for the articles they learn from and quote. If the publishers win, AI developers could owe damages for each article copied, and paying for content would become the norm rather than the exception.

The words the defendants used

The brief leans heavily on the defendants' own documents. Its opening pages attribute the "astonishing theft" line to Microsoft's Director of Applied Science, identified later in the filing as Dr. Brent Hecht. The brief pairs it with another document describing the practice as perhaps the "largest theft of labor in human history," and records the same Microsoft executive recognising that a fair use victory for the defendants would arguably "make a complete mockery of the idea of 'fair use.'" The fuller version of the theft passage, quoted in the section on data acquisition, adds that "almost no one intended for content they created to be used in this fashion, nor are they compensated for its use."

OpenAI's side of the record is quoted just as directly. Nick Turley, identified in the brief as OpenAI's Head of ChatGPT, wrote that publishers faced an "existential threat" from the company's products, which he said "are largely substitutive, period" and "will get more and more substitutive as they get better." According to the publishers, OpenAI understood that existential threat as early as June 2023. An OpenAI software engineer, the brief states, put the traffic problem more bluntly: "no matter how prominently we show the links, users won't click."

Greg Brockman, OpenAI's co-founder and president, appears repeatedly. The brief quotes him telling colleagues that "we are excellent at news btw. every time i do any generative stuff on NYT it seems to predict the next sentence pretty well." Another message quoted in the filing describes the model as "particularly good at predicting text of news articles like whenever i have it complete in the middle of a sentence in a NYT article, it seems to complete the sentence on point."

Satya Nadella, Microsoft's chief executive, is cited from his deposition. According to the brief, he agreed under oath that conversing with chatbots "has substituted ... giving you the information right there on the website on the AI platform versus needing to go to the underlying source." He also testified, the publishers say, that "anything that is paywalled should be licensed by anyone who wants to use it [whether] for ... grounding or training," and that he would have invoked Microsoft's right to require OpenAI to retrain its models had he known OpenAI scraped and trained on paywalled material.

Then there is the phrase the publishers return to most often. A Microsoft document, according to the brief, describes the company's position this way: "Our AI content strategy has started a 'doom loop' that will hurt the performance of our models and the entire web at the same time: It is highly unusual that an end-product threatens the economic foundations of its essential suppliers, but that is the situation we have created for our LLM business with respect to its 'content supply chain.'" Another Microsoft line cited in the filing is shorter: "LLMs are a product that destroys its supply chain."

Why would a defendant's own staff write this? The brief does not speculate. It simply places the passages next to the legal standard. Substitution, the publishers note, is what the Supreme Court in Warhol v. Goldsmith treated as the central concern of the first fair use factor.

Five stages of copying

The publishers divide the alleged infringement into five uses: acquisition, training, grounding, outputs and what both companies apparently called "horse trading." Claims on grounding and outputs are pressed against Microsoft only. According to a footnote, the issue is not ripe against OpenAI until the court resolves a pending sanctions motion.

Acquisition

The brief describes OpenAI building its own scrapes and downloading third-party archives. Beginning with GPT-2, OpenAI created a dataset called WebText, in which news articles were the most prevalent content type, according to the filing. An expanded WebText2 followed for GPT-3 and GPT-3.5 and contains at least 6,552 scraped works from The Times, 18,609 from the Daily News plaintiffs and 66,780 from Ziff Davis, the brief states. A separate table for Ziff Davis lists 132,820 copies under a row labelled WebText rather than WebText2; the brief does not reconcile the two figures.

From GPT-3 onward, OpenAI also drew on Common Crawl, a nonprofit archive the publishers say has received funding from the defendants and that they never authorised to scrape their sites. One OpenAI Slack discussion quoted in the brief listed the top 100 domains by document count in a filtered Common Crawl training set. That list included 2,182,079 articles from chicagotribune.com and 2,064,805 from nytimes.com. According to the publishers, OpenAI downloaded the Common Crawl datasets repeatedly over several years and planned to keep the files "forever" so it could recreate them later. Reporting in November 2025 documented how Common Crawl captures paywalled articles because its scraper never runs the subscription-check code, and the News/Media Alliance demanded in April 2026 that the archive stop enabling unauthorised use.

By drawing on third parties and anonymous scrapes, the brief argues, OpenAI sidestepped the robots.txt regime that publishers use to restrict crawling. OpenAI's corporate representative testified he was unaware of "any effort to detect paywall content in its training datasets" or "to remove paywall content from its training datasets," according to the filing. The company's general practice when crawling directly "did not include reviewing websites['] ... Terms of Use or Service."

A footnote adds a detail about opt-outs. In 2024 OpenAI said it would release a tool called Media Manager to let content owners express preferences about AI use of their material, including opting out of scraping. According to the brief, OpenAI has since abandoned the project.

One acquisition carries a licence dispute of its own. OpenAI obtained the New York Times Annotated Corpus from the Linguistic Data Consortium, a set of over 1.8 million articles covering nearly everything the paper published between January 1, 1987 and June 19, 2007. The licence limited use to "non-commercial linguistic education, research and technology development." According to the brief, OpenAI employees knew it "would not be appropriate" to use the corpus "to train a model," and did so anyway. The Times seeks summary judgment on 928,872 unique works drawn from it.

Paywalls

The paywall material is among the most pointed in the filing. According to the brief, when OpenAI's Nick Ryder told Brockman about "a hack to get around nytimes paywall" to help with Brockman's effort to scrape The Times's site, Brockman replied "ah nice." OpenAI researchers explained in March 2023 that ChatGPT was able to "access the content" behind a paywall, the publishers say.

The same concern extends to the GPT Store. The brief lists custom GPTs hosted on OpenAI's platform with names including "Bypass Paywall," "Remove Paywall," "NYTimesGPT" and "News Summarizer Ace." One, called "Article Reader," described itself as a tool to "summarize and extract key points from paid (paywalled items). Just submit the URL." In one exchange cited by the publishers, ChatGPT invoked a custom GPT called WebPilot, which OpenAI itself had identified as purporting to interact with content "that may be subject to paywalls or access controls," to retrieve a Times article and output its text verbatim.

Training and memorization

The brief describes three phases of model training, each involving further copies. During pre-training, OpenAI passed millions of copies of the publishers' articles through its models, according to the filing. During mid-training, curated datasets targeting news were introduced late in the training cycle. OpenAI's stated goal, per the brief, was to "crush freshness" in "the domain of real-world news," and internal material described "current events / news articles (NYT data)" as an area to improve. The mid-training datasets contain over 91,692 copies of the publishers' works.

Post-training, the filing says, taught models to formulate search queries, retrieve web content and ground answers on it, with training data instructing models to retrieve and copy from the publishers' sites.

The evidence on memorization and regurgitation follows a timeline. By November 2019, according to the brief, OpenAI was internally concerned about models "accidentally regenerating copyrighted works." By 2020 it understood that even a single exposure during training "causes some degree of memorization." In June 2022, employees said GPT-4 would have "memorized a ton of data and therefore will be insanely good at regurgitation." OpenAI's vice president of research is quoted as saying: "We train our networks to memorize the training data - that's their objective."

The filing also records a detail from the period before ChatGPT's release. Dario Amodei, then a senior researcher at OpenAI and now chief executive of Anthropic, listed "News Generation" as one of GPT-3's skills in a presentation, with "What's the NYT saying today?" as a sample query for the publishing category, according to the brief.

The Bloom filter

Immediately after the lawsuits were filed, OpenAI searched for and copied the publishers' content it considered most likely to be output by its models, and used it to populate a Bloom filter designed to suppress that output, the brief states. According to the publishers, OpenAI did not suppress output of content from any entity that had not sued it. They argue the purpose was to stop plaintiffs gathering evidence rather than to prevent infringement. Hecht, the Microsoft scientist, called the approach an "accidental cover up," the brief says, because it would leave "people who have a right over the content having less visibility into what was used for training."

The filter itself is counted as infringement. The publishers list 140,097 Times copies and 270,250 Daily News copies in what the brief calls OpenAI's "Giraffe Bloom Filter."

Grounding in Copilot

Grounding is the process by which a chatbot fetches live content to answer a query. Microsoft's own description, quoted in the brief, has a user query trigger a retrieval function; the retrieved content is merged into the model's context window along with the prompt, and the model generates an answer from both.

According to the publishers, at least two million Copilot conversations copied and grounded outputs on their website content. Those conversations include 34,970 unique URLs for Times asserted works and 35,202 for the Daily News plaintiffs. Microsoft's technical expert, Dr. John Lafferty, conceded at least 88,178 instances in which Microsoft copied at least six non-overlapping 16-word sequences, or 96 words, from a plaintiff work during retrieval, the brief states. One retrieval grounded on at least 1,820 words of a single article.

The examples are specific. In one conversation, Copilot retrieved more than 340 words from the Times article "Trump and Allies Forge Plans to Increase Presidential Power in 2025" and output around 270 of them to the user. In another, it retrieved around 340 words from a Denver Post investigation into remains of 230 Native Americans held in Colorado museums and returned nearly all of that text. A Mother Jones request produced 185 verbatim words. The Times and the Daily News plaintiffs seek judgment on samples of 58 and 68 works respectively, reserving the rest for trial.

Blocking did not stop the retrieval, according to the filing. OpenAI created blocklists for ChatGPT grounding but did not block any plaintiff website until September 2023, and blocked others as late as January 2025. Even after the publishers followed the opt-out instructions the defendants published, both companies continued to access their sites for grounding, the brief states.

Horse trading

The fifth category, redacted in the earlier public version, turns out to describe data exchanges between the two defendants. The brief says the companies used publisher content "as a form of currency through self-described 'horse-trading' deals," selling it to each other on at least three occasions.

The first ran from OpenAI to Microsoft. In September 2020, OpenAI delivered its GPT-3 training data to Microsoft, which accessed it in 2021 to evaluate how to build OpenAI's models into its own products, according to the brief. GPT-3 had already been trained by then, the publishers note, so the transfer bore no relationship to training any model at issue.

The second, codenamed Project Taxi, ran the other way. Over three years from 2019 to 2022, Microsoft provided OpenAI with a copy of the Bing Index, which Microsoft's counsel has described as a compilation of "billions" of webpages gathered for conventional search. OpenAI convinced Microsoft to "sell" it the index, per the brief, for a price that remains redacted. Bob McGrew, OpenAI's chief research officer, described the arrangement in October 2021: "the idea [was] that we're giving data to them and they are giving data to us." The publishers say neither company asked copyright holders for permission.

The third, Project Mango, was a joint crawl. Microsoft developed and operated the "Mango crawler" on OpenAI's behalf to "collect as many of the documents as possible for their training," according to the brief, and OpenAI paid Microsoft an amount the filing still redacts. The resulting training dataset contains copies of at least 160,903 unique works from the plaintiffs. The publishers argue this also breached industry norms, because they had allowed Microsoft to crawl for search, not to supply training data to a partner.

The count

The Times and the Daily News plaintiffs seek summary judgment only on works that clear a matching threshold: 80 percent of three-word sequences contained in the defendants' copy, plus a shared 16-word sequence, after quotations and non-letter characters are removed. The comparisons, prepared under the method of their expert Dr. Tom Goldstein, were lodged with the court on a hard drive. The resulting figures are large.

Infringing actThe TimesDaily News plaintiffs
Training3,924,653 copies (944,655 unique works)7,387,394 copies (1,388,384 unique works)
OpenAI's WebText2 scraping12,714 copies (6,552 unique works)37,264 copies (18,609 unique works)
Project Mango scraping28,744 copies (27,453 unique works)34,169 copies (32,087 unique works)
Common Crawl downloads3,551,481 copies (930,825 unique works)6,940,478 copies (1,368,704 unique works)
GPT-3 training data sent to Microsoft556,185 copies (365,928 unique works)531,716 copies (397,150 unique works)
Giraffe Bloom filter140,097 copies (128,266 unique works)270,250 copies (207,152 unique works)

Ziff Davis applies a stricter test, counting only exact matches or revised versions with at least 80 percent directional word containment and one-gram Jaccard similarity above 50 percent. On that basis it claims 6,271,629 training copies covering 1,636,586 unique works, 5,950,361 Common Crawl copies and 1,052,721 copies in the GPT-3 data sent to Microsoft. The Center for Investigative Reporting seeks judgment on 682 registered works for training and 743 for acquisition, while it and The Intercept point to tens of thousands and thousands of unregistered works respectively for their separate DMCA claims.

The DMCA count concerns copyright management information: titles, bylines, copyright notices and links to terms of use. The publishers want partial summary judgment on three of the four elements of a claim under Section 1202(b)(1), leaving knowledge for trial.

The brief's figures are almost total. According to the filing, The Times's copyright notice is absent from 99.9 percent of copies in OpenAI's training sets, Ziff Davis's from 99.8 percent, Mother Jones's from 99.99 percent and The Intercept's from every copy. For the Daily News plaintiffs, the notice is missing in 100 percent of cases for five titles and at least 99.4 percent for the rest. In absolute terms, OpenAI removed at least one piece of such information from 3,002,796 copies of Times works, 2,638,497 Daily News copies and 3,462,149 Ziff Davis copies, along with 370,993 pieces from 101,092 CIR copies and 46,142 pieces from 13,353 Intercept copies.

The mechanism, per the brief, was a set of text extractors the filing names as Dragnet, Newspaper and Gutentag. These targeted terms such as "byline," "copyright" and the copyright symbol where they appeared as page furniture. In articles about copyright itself, the extractors removed the notice but kept the word "copyright" in the body text, the publishers say. OpenAI employees admitted removing text the company "wouldn't want the model to be outputting," including "copyright notices," according to the brief, and that terms-of-service links "would not be something that [OpenAI] would want inputted into the model."

The traffic evidence

The most commercially significant numbers concern clicks. Microsoft's representative data, according to the brief, shows the overall click-through rate reduction between Bing Chat and Bing web search was 87 to 93 percent for The Times's sites, 83 to 91 percent for the Daily News sites and 51 to 94 percent for Ziff Davis sites. The earlier public version redacted these figures.

Microsoft marketed exactly that behaviour, the publishers argue. Copilot's home page greeted users with: "Instead of clicking through links, we can talk through whatever you're curious about." Elsewhere the brief quotes the same page with "talk about" in place of "talk through," without explaining the difference. Microsoft's lead counsel, per the filing, described the advantage of chatbots as being "actually designed to answer your question" rather than "just giving you a series of links." Turley, the brief says, wrote that once ChatGPT's browse function gave an answer there was "no good reason to click."

Third-party figures reinforce the argument. A study cited in the brief found 87.78 percent of ChatGPT users visit no external website during a search, against 26.91 percent of Google users. Microsoft's own survey expert reported that 37.2 percent of Copilot respondents would otherwise use conventional web search. A study by multiple OpenAI authors found that seeking information, including current events, accounted for 24 percent of ChatGPT usage in 2025, and the publishers put ChatGPT at 900 million weekly users by February 2026.

The subscriber survey is perhaps the sharpest single statistic. Among Times subscribers who used or paid for ChatGPT for news, 36 percent said their use "means I no longer need news from the New York Times at all," according to the brief.

The filing also cites Cloudflare's chief executive for the ratio of pages scraped to visitors sent: 2 to 1 for Google in 2015, and by June 2025 18 to 1 for Google and 1,500 to 1 for OpenAI. Cloudflare's chief executive first described that 1,500 figure in May 2025, and per-operator data published in August 2025 put OpenAI at 1,091 crawls per referral.

Google enters the argument

Google is not a defendant, but its products run through the market-harm section. The brief argues that ChatGPT's release in November 2022 pushed other companies to fast-track rival products, most significantly Google's AI Overviews. According to the publishers, the defendants' own experts concede harm from that feature. OpenAI's economist Dr. Avi Goldfarb opined that "Google's introduction of AI overviews may have depressed search referrals by 20 to 60 percent" for the Daily News plaintiffs. OpenAI's media expert, Dr. Sinnreich, cited a 2026 Reuters Institute analysis finding referrals via Google Discover fell from over 5 billion a month to fewer than 4 billion, and via Search from well over 3 billion to slightly more than 2 billion, since AI Overviews launched.

The publishers use this to make a structural point. Each AI company, they argue, is better off taking content for free while others pay, a prisoners' dilemma that a ruling against fair use would resolve by placing every developer on the same footing. Separate research has put a causal figure on the Google side of that dynamic: a randomised experiment with 1,065 users found AI Overviews cut outbound publisher clicks by 39.8 percent.

Synthetic news at $6,800 per million articles

The filing devotes a section to pink slime, which it defines as unreliable, low-quality content often plagiarised or remixed from legitimate sources. Using OpenAI's published API pricing, the brief calculates that one million news-style articles of 500 words each would cost roughly $6,800 to generate, with no author, editor or reporter. Its example is Prism News, which operated 200 AI-generated publications posing as local newsrooms and hobby sites with four employees, according to the publishers. The brief says the defendants knew their products would spread such material and that it was often indistinguishable to readers from original reporting.

The publishers also note that chatbots invent content and attribute it to them. The Copilot and ChatGPT conversations produced in discovery include hallucinated material "that plausibly but falsely purports to originate from a Plaintiff," the brief states.

A licensing market the defendants helped build

On the fourth fair use factor, the publishers argue that a market for licensing news to AI developers already exists and that OpenAI and Microsoft participate in it. OpenAI has signed Media Publisher Agreements covering historical archives and ongoing content, and Microsoft has signed comparable deals, though the number of agreements and the payment ranges are redacted. Amazon has entered at least 23 Data Service Agreements, per the brief, and Google, Meta, Perplexity, ProRata.ai and Mistral AI have signed deals with news publishers. The filing lists intermediaries including TollBit, Cloudflare, ScalePost, Human Native AI and the Copyright Clearance Center.

Even Microsoft's economic expert, Dr. Rao, used real-world licences for training and grounding to inform his damages calculations, the publishers say. That, they argue, satisfies the Second Circuit's requirement that a market be "traditional, reasonable, or likely to be developed."

What the publishers want

The requested relief comes in five parts. The publishers ask the court to find prima facie infringement for acquiring, training on and trading their works, and for Microsoft's grounding and outputs; to reject fair use for all of those uses; to find the first three elements of the DMCA claim established; to rule that statutory damages, if awarded, run per article rather than per newspaper issue; and to dispose of the defendants' remaining affirmative defences for lack of evidence.

The per-article point carries the most financial weight. The publishers rely on the Second Circuit's 2016 decision in EMI Christian Music Group v. MP3tunes, which allowed separate awards for songs released as singles even when they also appeared on albums. Articles have been available online individually since at least the late 1990s, a point OpenAI's own expert conceded, according to the brief. With asserted works numbering in the millions, the choice between per-issue and per-article counting determines how many separate statutory awards are available at all.

The brief also records that OpenAI will offer no evidence it believed at the time that its conduct was justified by fair use, citing a February 6, 2026 order in the parallel class case and an email in which OpenAI's counsel agreed the same applies to the news cases.

The defendants' account

The September 17 filing is advocacy, and the two companies tell a very different story. In their own September 4 motions, OpenAI reported that its expert found 24 instances of verbatim regurgitation in a sample of 20 million ChatGPT conversation logs, a rate of 0.00012 percent, and argued that its models learned from a broad, undifferentiated sweep of internet text. OpenAI also contended that copies made through its browsing agent before publishers updated robots.txt were impliedly licensed, and that its text extractors shed copyright notices as a side effect of cleaning markup rather than by design. Microsoft moved for judgment on all claims against it.

The publishers' brief anticipates part of that. It quotes OpenAI expert Dr. Chris Callison-Burch stating that "any individual source" is "entirely fungible" for training, and Lafferty stating that removing any single work "would not materially alter the statistical structure the model learns." The publishers turn this around: if the content was unnecessary to build general-purpose models, they argue, copying it in full cannot be justified under the third fair use factor. The United States Department of Justice filed a statement supporting OpenAI's fair use defence on September 3, 2026.

Precedent points both ways. The brief leans on Judge Vince Chhabria's 2025 ruling in Kadrey v. Meta, which found for Meta on its record but observed that news publishers could present "even stronger arguments against fair use" than authors. Anthropic, meanwhile, agreed to a $1.5 billion settlement in a separate authors' case in September 2025.

Why it matters for marketing and media

For publishers, the fourth-factor section describes the economics PPC Land has tracked for two years, now with the defendants' internal data attached. The Times has sued three AI companies while licensing to a fourth, and its chief executive disclosed in August 2026 that the company spends close to $2 billion a year on journalism. A ruling that training and grounding are fair use removes the bargaining position behind those negotiations. A ruling the other way, with per-article statutory damages, puts a floor under every licence in the sector.

For search and content teams, the Copilot click-through ranges are the first figures in this litigation drawn from a defendant's own measurement of how an answer engine performs against a results page for the same domains. They sit alongside Chartbeat data showing small publishers lost 60 percent of search traffic over two years. The blocklist dates matter too. That both defendants allegedly kept retrieving pages after publishers followed their published opt-out instructions bears directly on how much control crawler configuration actually provides.

For media buyers, the $6,800 figure prices synthetic inventory. Integral Ad Science flagged AI-generated slop sites as a quality threat in July 2025, and the same programmatic budgets that reach licensed newsrooms can reach sites produced at a fraction of a cent per article.

The commercial ties between the defendants have also shifted since much of this conduct occurred. The brief describes a 20 percent revenue share and more than $10 billion in Microsoft funding, but the two companies amended their agreement on April 27, 2026, dropping some revenue share payments and making the intellectual property licence non-exclusive.

No hearing date has been set. Both defendants have requested oral argument, and a sanctions motion over discovery remains pending, which the publishers say must be resolved before grounding and output claims against OpenAI can be decided.

Timeline

Summary

Who: The New York Times Company, the Daily News plaintiffs (New York Daily News, Chicago Tribune, Orlando Sentinel, Sun-Sentinel, Denver Post, Pioneer Press, Orange County Register and Mercury News), Ziff Davis, the Center for Investigative Reporting and The Intercept, against OpenAI and Microsoft, before Judge Sidney H. Stein.

What: The public version of the publishers' 92-page combined summary judgment brief, which sets out internal documents in which Microsoft staff described AI data practices as "an astonishing theft," OpenAI's head of ChatGPT called the products "largely substitutive," and Microsoft data recorded click-through rate reductions of up to 94 percent for Copilot against Bing search. It also details paywall circumvention, the Project Taxi and Project Mango data exchanges, a Bloom filter built after the suits were filed, millions of training copies, and copyright notice removal in 99.4 to 100 percent of copies.

When: Filed on September 17, 2026, following cross-motions for summary judgment lodged on September 4, 2026. The conduct described runs from 2017 to 2025.

Where: The United States District Court for the Southern District of New York, in the multidistrict litigation numbered 25-md-3143.

Why: The publishers seek rulings that the defendants infringed at five stages, that fair use does not apply, that three DMCA elements are met and that statutory damages run per article. The outcome will shape whether AI developers must license news content and on what terms, with consequences for publisher traffic, search visibility and the quality of the inventory advertisers buy.