Three sets of lawyers filed for summary judgment on the same Friday, and the two records they describe barely overlap. OpenAI counted 24 verbatim reproductions in 20 million chat logs. The publishers counted 10.8 million articles taken.

OpenAI, Microsoft and a group of five news organisations each moved for summary judgment on September 4, 2026 in the consolidated copyright proceedings before Judge Sidney H. Stein in the Southern District of New York. The filings, docketed in the multidistrict litigation numbered 25-md-3143, ask the court to resolve central questions of liability and fair use before any jury is empanelled. Redactions cover much of what each side considers its strongest material, and the public versions arrived with black bars over traffic figures, revenue comparisons and internal executive commentary.

The three briefs describe the same conduct in incompatible terms. According to OpenAI, its models were trained on a broad and undifferentiated sweep of internet text, they almost never reproduce the articles they read, and the publishers suing it have grown rather than shrunk since ChatGPT launched. According to the news plaintiffs, the defendants copied entire archives at five separate stages of an artificial intelligence pipeline, built commercial products that answer the questions readers used to bring to news sites, and are now asking a court to call the whole arrangement fair.

Who is in the case

The plaintiff group is not homogeneous. It comprises The New York Times Company, which the brief describes as employing 6,000 staff including more than 3,000 in the newsroom and holding 149 Pulitzer Prizes; the Daily News Plaintiffs, publishers of the New York Daily News, Chicago Tribune, Orlando Sentinel, Sun-Sentinel, Denver Post, Pioneer Press, Orange County Register and Mercury News, with 44 Pulitzer Prizes between them; Ziff Davis, whose properties include IGN, Everyday Health, CNET, ZDNET, PCMag and Mashable; the Center for Investigative Reporting, which operates Mother Jones and Reveal; and The Intercept.

Ziff Davis sued OpenAI in April 2025 over content drawn from more than 45 properties, joining actions that The Times had begun in December 2023. The cases were later consolidated.

Together the plaintiffs have identified more than 10.8 million asserted works, according to OpenAI's brief. Of those, OpenAI says the plaintiffs' own experts contend that only 6.3 million were used in training, leaving roughly 4.5 million for which OpenAI argues no evidence of copying exists at all. The Times alone asserts 6,030,928 articles, with publication dates spanning 1928 to 2024.

The regurgitation numbers

OpenAI's memorandum, certified at 8,185 words, builds its factual case on how rarely its models reproduce text.

When OpenAI's expert searched a sample of 20 million ChatGPT conversation logs produced in discovery, according to the brief, he found 24 instances of verbatim regurgitation, a rate of 0.00012 percent. The plaintiffs' own experts produced figures in the same range: 0.00011 percent for The Times and the Daily News Plaintiffs, 0.00002 percent for Ziff Davis, 0.000008 percent for the Center for Investigative Reporting and 0.000002 percent for The Intercept. The longest verbatim passages identified ran to 29 words and 43 words.

Those logs exist because of an earlier discovery fight. A magistrate judge ordered OpenAI to produce 20 million anonymised conversations in November 2025, and a stay request was denied on November 13 of that year.

Two further studies appear in the OpenAI filing. In the first, prompts about 710 randomly selected asserted works produced more than 68,000 model outputs with no instances of regurgitation. In the second, adversarial prompting aimed at 19 works, one from each of 19 outlets across the five plaintiffs, used more than 9,400 prompts and extracted an average of 2.3 percent of the content of the targeted works.

The plaintiffs ran larger experiments. According to OpenAI, their experts used more than 250 million prompts, non-public GPT models and configurations available only through the application programming interface rather than the consumer chatbot. Even so, the brief states, those attempts triggered partial regurgitation of fewer than 2.6 percent of the asserted works, and within that subset extracted an average of under 4.4 percent of each targeted work. OpenAI sets that figure against the 16 percent that plaintiffs' experts extracted in the Google Books litigation, where the Second Circuit found fair use.

The age of the corpus does further work in the argument. More than 95 percent of the Times works at issue were published before OpenAI was founded in November 2015, according to the brief, and two-thirds were published before 1981. Among the works OpenAI singles out as carrying thin copyright protection are a set of sports results from July 13, 1941, a television schedule for the 1976 Winter Olympics and Colorado snow totals for October 19, 2023.

Browse, robots.txt and the implied licence

A separate section of the OpenAI motion addresses Browse, the web-browsing function inside ChatGPT that the plaintiffs describe as retrieval-augmented generation.

OpenAI implemented the feature in early 2023 to address what the brief calls a knowledge cutoff, after which the model was blind to the world. The mechanism it describes is sequential: a user question becomes a search query, the query goes to Microsoft Bing, Bing returns links and short snippets, ChatGPT decides which pages might be relevant, and only then does it request the page content itself. Facts extracted from those pages inform a synthesised answer with links back to sources. The brief reproduces a real user exchange from September 26, 2024 about when Hurricane Helene was expected to make landfall.

More than 96 percent of the instances in which the plaintiffs claim Browse made an infringing copy involved only Bing search-result snippets, according to OpenAI, with no request sent to the underlying website. The index that returns those snippets is built by bingbot, Microsoft's crawler.

The opt-out argument rests on timing. OpenAI disclosed the ChatGPT-User agent in a March 2023 blog post and stated it was configured to honour robots.txt files. It added OAI-SearchBot for Browse in 2024. Its later training crawler, GPTBot, is a separate agent from those at issue here. A two-line robots.txt entry would have refused the browsing agent, and OpenAI also published the IP addresses used by its bots.

The brief then records what the publishers did. Times employees discussed the announcement on April 1, 2023, nine days after it was made. The Times waited nearly a year before blocking OpenAI's user agent and implemented a hard block for ChatGPT-User in April 2024. The Intercept blocked in January 2024. Dates for the other plaintiffs are redacted. On that record, OpenAI argues, any copies made before the blocks went up were impliedly licensed under the line of cases running from Field v. Google.

A footnote cites Reuters Institute research from February 22, 2024 finding that by the end of 2023, 48 percent of the most widely used news websites across ten countries were blocking OpenAI's crawlers.

The publishers also press a claim under Section 1202(b)(1) of the Digital Millennium Copyright Act, which prohibits the intentional removal of copyright management information.

OpenAI's answer is mechanical. Raw web data, the brief states, consists of thousands of lines of markup and non-language content that would produce a model generating markup if used directly. Text-extraction scripts exist to pull usable natural language out of that mess, and any copyright notice left behind is a side effect rather than a target. The company cites the Second Circuit's 2026 decision in McGucken v. Shutterstock and an August 2026 decision in Beaulier v. Meta Platforms for the proposition that uniform transformation that incidentally sheds such information does not establish intent.

On concealment, OpenAI's expert reviewed all 68 alleged regurgitations identified in the parties' opening expert reports, drawn from 108 million conversations, and concluded that the source of the text would have been obvious to a user in 63 of them, or 92.7 percent.

The claim has a difficult history in this district. A federal court dismissed a comparable DMCA case brought by Raw Story and AlterNet in November 2024 on standing grounds.

What the publishers filed

The plaintiffs' 92-page combined brief, signed by Ian Crosby of Susman Godfrey, seeks judgment on liability for copying at five stages: acquisition, training, grounding, outputs, and a fifth category whose description is redacted throughout. Grounding and output claims are pressed against Microsoft only, because the brief states the issue is not ripe against OpenAI until the court resolves a pending sanctions motion.

The acquisition section runs through datasets. OpenAI created WebText for GPT-2 and an expanded WebText2 for GPT-3 and GPT-3.5. From GPT-3 onward it also drew on Common Crawl, whose archives the plaintiffs say they never authorised and whose own terms require users to observe the terms of the sites it scrapes. News publishers demanded in April 2026 that Common Crawl stop enabling unauthorised use of member content, and reporting in November 2025 documented how the archive captures paywalled articles because its scraper never runs the subscription-check code.

One acquisition is described without redaction. OpenAI obtained the New York Times Annotated Corpus from the Linguistic Data Consortium, a dataset containing nearly every article published by the paper between January 1, 1987 and June 19, 2007, more than 1.8 million articles. The user licence limited use to non-commercial linguistic education, research and technology development.

The brief also lists third-party tools built on OpenAI's platform, including custom GPTs named News Summarizer Ace, NYTimesGPT, Bypass Paywall, Remove Paywall and Article Reader, the last described as summarising and extracting key points from paywalled items given a URL.

Rather than assert every work, the plaintiffs apply matching thresholds. The Times and the Daily News Plaintiffs seek judgment only for works where 80 percent of three-word sequences appear in the defendants' copy and the two texts share a 16-word sequence, with quoted material and non-letter characters stripped before comparison. Ziff Davis uses a two-tier test: an exact match where directional word containment reaches 95 percent, and a revised version where containment reaches 80 percent with one-gram Jaccard similarity above 50 percent. Everything below those thresholds goes to a jury.

The traffic argument

Both sides claim the same terrain on the fourth fair use factor, and both have hired economists.

According to OpenAI, Avi Goldfarb analysed individual user data and found no impact from ChatGPT adoption on visits to the plaintiffs' websites. The brief notes that Times revenues have risen, that the company surpassed its goal of 12 million subscribers in 2025 and is on pace for 15 million by the end of 2027, and that Ziff Davis recorded 3.5 percent revenue growth in 2025 while declining to name generative artificial intelligence as a material risk in its 2025 annual report.

Those markers have moved since. The Times reported second-quarter digital advertising revenue of $114.0 million, up 20.7 percent, on August 5, 2026. The following day Ziff Davis recorded a $54.8 million goodwill impairment on its health media unit as advertising revenue fell.

The publishers argue the question is legally beside the point, and that fair use is an affirmative defence carrying the burden of proving no harm to existing or potential markets. Their own evidence is largely redacted. What survives includes Microsoft's internal recording of click-through rate drops for Copilot compared with conventional Bing search across the Times, Daily News and Ziff Davis domains, with the figures blacked out, and a passage from Microsoft's lead counsel describing the advantage of chatbots as being designed to answer a question rather than return links to look through.

One external figure appears in full. Citing Cloudflare's chief executive, the brief states that in 2015 Google scraped two pages for every visitor it sent back, and that by June 2025 the ratio stood at 18 to 1 for Google and 1,500 to 1 for OpenAI. Cloudflare has published a series of crawl-to-referral disclosures since May 2025, including per-operator figures showing OpenAI at 1,091 crawls per referral in mid-2025 and a dashboard opened in July 2026 recording ratios from 118 to nearly 50,000.

Market dilution and synthetic news

The plaintiffs devote a section to what they call pink slime, low-quality material either plagiarised or remixed from legitimate sources and published under the appearance of local journalism.

The cost argument is specific. Using OpenAI's published API pricing, the brief calculates that generating one million news-style articles averaging 500 words each would cost roughly $6,800 and require no author, editor or reporter. It offers Prism News as an example, describing an operation running 200 artificial intelligence-generated publications posing as local newsrooms and hobby sites with four employees.

Advertising measurement firms have tracked the same supply. Integral Ad Science identified machine-generated sites as a quality threat in July 2025, citing forecasts that as much as 90 percent of web content could be machine-generated by 2026 and individual sites producing up to 1,200 articles a day.

The legal theory attached to this is market dilution, drawn from Judge Vince Chhabria's ruling in Kadrey v. Meta, which the plaintiffs quote for the proposition that markets for news articles may be more vulnerable to indirect competition from generated outputs than other categories.

Microsoft's separate motion

Microsoft filed its own notice of motion the same day, seeking judgment on all claims against it in the Times's third amended complaint and the Daily News Plaintiffs' first amended complaint.

The supporting record is unusually broad. It includes declarations from John D. Lafferty on technical questions, Catherine Tucker and On Amir on economic and consumer issues, Christopher A. Bail, and company witnesses including Sarah Bird, Jordan Usdan, Padma Gaggara, Elbio Abib, Kate Cook and Fabrice Canel, the product manager most publicly associated with Microsoft's crawler until his retirement in July 2026. Annette L. Hurst of Orrick, Herrington and Sutcliffe signed for Microsoft.

The commercial relationship underlying the joint defence has itself changed during the litigation. Microsoft and OpenAI amended their agreement on April 27, 2026, dropping revenue share payments and converting the intellectual property licence from exclusive to non-exclusive.

Jason Kint, chief executive of Digital Content Next, posted a public reading of the filings within hours. He described reading the defence arguments first and ending where he started, and said he was not impressed by OpenAI's motion, summarising its position as an argument that Times traffic is up, that the plaintiffs themselves use OpenAI's models, and that what the models extract is facts rather than protected expression.

Why this matters for the marketing industry

The dispute is not confined to two newsrooms and two technology companies. It sets the terms on which content acquired without payment can be used to build products that compete for the same attention, and the marketing supply chain sits downstream of that answer.

For publishers, the fourth-factor argument determines whether a licensing market exists at all. The Times has sued three artificial intelligence companies while licensing to a fourth, and its chief executive disclosed in August 2026 that the company spends close to $2 billion a year producing about half a million works. A ruling that training and grounding are fair use removes the bargaining position behind every one of those negotiations. A ruling the other way establishes a price floor across an industry that has so far set values in private.

For anyone buying media, the pink slime section describes the inventory problem in cost terms rather than quality terms. Material that costs $6,800 per million articles to produce competes for the same programmatic budgets as material that costs thousands of dollars per piece.

For search and content teams, the robots.txt record in OpenAI's brief is the operative detail. The argument is that failing to update a text file for a year amounted to consent. Whether or not the court accepts it, the filing establishes that crawler configuration is now litigation evidence.

The wider legal picture remains unsettled. Anthropic agreed to a $1.5 billion settlement in a separate authors' case in September 2025, the largest publicly reported figure attached to a copyright claim in the sector, while Meta won summary judgment on fair use in June 2025 on a record its judge described as thin. The Department of Justice filed a statement supporting OpenAI's fair use defence on September 3, 2026, the day before these motions landed.

Oral argument has been requested by both defendants. No date has been set.

Timeline

Summary

Who: OpenAI and Microsoft on one side, and on the other The New York Times Company, the Daily News Plaintiffs, Ziff Davis, the Center for Investigative Reporting and The Intercept. Judge Sidney H. Stein presides in the Southern District of New York, with Magistrate Judge Ona T. Wang handling discovery.

What: Cross-motions for summary judgment. OpenAI seeks a ruling that pretraining and the Browse function are fair use, that Browse copies made before publishers updated robots.txt were impliedly licensed, and that the DMCA claim fails as a matter of law. Microsoft moves for judgment on all claims against it. The publishers seek judgment on liability for copying at five stages, rejection of the fair use defence, partial judgment on three elements of the DMCA claim, and a ruling that statutory damages run per article rather than per issue.

When: All three filings are dated September 4, 2026. The underlying conduct dates from 2019 onward, and the asserted works span 1928 to 2024.

Where: The multidistrict litigation numbered 25-md-3143 in the United States District Court for the Southern District of New York, at 500 Pearl Street in Manhattan. Both defendants have requested oral argument.

Why: The motions force the court to choose between two accounts of the same evidence. OpenAI's record shows verbatim reproduction at 0.00012 percent of sampled conversations and publisher revenues rising. The publishers' record shows 10.8 million articles copied, click-through rates falling inside Copilot relative to Bing search, and a licensing market that other artificial intelligence developers have already entered. Whichever account prevails sets the default price of news content for every generative system built on the open web.