A three-judge panel of the United States Court of Appeals for the Ninth Circuit on September 16, 2026 affirmed the dismissal of a Digital Millennium Copyright Act claim brought by anonymous programmers against GitHub, Microsoft and OpenAI, holding that the AI coding tools Copilot and Codex generate new works rather than stripping attribution from existing ones. The opinion arrived 15 days after the Department of Justice told a New York federal court that training large language models on copyrighted text is fair use.
In Short
Some programmers said that when an AI coding assistant spits out code resembling theirs without their name or licence attached, the companies behind it broke a law against deleting copyright notices. The appeals court said no, because the tool writes something new rather than taking a copy and scrubbing the name off it. That makes one of the harshest penalty routes in US copyright law much harder to use against AI companies, while leaving ordinary copyright claims and questions about training data untouched.
What the Ninth Circuit decided
The case is Doe v. GitHub, Inc., docket number 24-7700, on appeal from the Northern District of California, where District Judge Jon S. Tigar presided under docket 4:22-cv-06823-JST. Circuit Judges Sidney R. Thomas and Eric D. Miller heard it with Stanley Blumenfeld, Jr., a district judge from the Central District of California sitting by designation. Argument took place in San Francisco on February 11, 2026. Judge Miller wrote the opinion, which runs to 18 pages and ends with a single word: affirmed.
The dispute concerns Section 1202(b) of the Digital Millennium Copyright Act (DMCA), a 1998 provision that makes it unlawful to intentionally remove or alter copyright management information, known as CMI. According to the opinion, the statute defines CMI to include the title of a work, the name of the author or copyright owner, and the "[t]erms and conditions for use of the work." The statute does not require anyone to attach CMI to a work. If CMI is present, however, deliberately removing or altering it is prohibited.
The plaintiffs are programmers who published copyrighted code on public GitHub repositories under open sourcelicences. Much of that code, according to the court, is released under terms permitting reuse on conditions, and "[o]ne of the most common conditions is attribution," meaning a copy of the licence, the author's name and the copyright notice must travel with any copy or derivative. Their complaint alleged that Copilot sometimes emits "essentially verbatim" copies of their code without those notices.
The court described the defendants' products in some detail. Microsoft owns GitHub, which the opinion calls the world's largest hosting service for open-source software, and is also a major investor in OpenAI. Codex is OpenAI's generative tool that drafts code autonomously from a programmer's prompts. Copilot is a paid subscription service developed jointly by GitHub and OpenAI that runs a modified version of Codex. According to the complaint as summarised by the court, Copilot was trained on billions of lines of publicly available code, including code from "all available public GitHub repositories." Because the two tools function similarly in every respect relevant to the case, the panel referred only to Copilot.
After two rounds of dismissals and amendments, the complaint had been reduced to three claims: one under the DMCA and two for breach of contract. Judge Tigar dismissed the DMCA claim with prejudice under Federal Rule of Civil Procedure 12(b)(6), reasoning that Section 1202(b) requires the copies to be identical, and certified the question for interlocutory appeal under 28 U.S.C. Section 1292(b). The contract claims survived a motion to dismiss and, according to the opinion, remain pending in the district court.
Two theories, one forfeited
The panel identified two distinct theories in the plaintiffs' case. The first, which Judge Miller labelled the input theory, held that the defendants violated the DMCA at the training stage by removing CMI from code before feeding it into Copilot. The second, the output theory, held that Copilot violates the statute whenever it returns memorised training data to users without the original notices.
The input theory never reached the merits. According to the opinion, the district court had asked at an early hearing whether copying training data into Copilot breached the attribution requirement of open-source licences, and plaintiffs' counsel replied: "Perhaps it doesn't." Judge Tigar then observed that the "complaint is not about training. It just isn't." His later written order stated that "Plaintiffs do not allege they were injured by Defendants' use of licensed code as training data." Plaintiffs did not contest that characterisation orally or in writing, and the appeals court treated the theory as forfeited.
That procedural finding matters more than its length suggests. The court did not decide whether stripping attribution before training could violate the DMCA. It simply declined to consider an argument the plaintiffs had not preserved.
Standing survived
The defendants had argued that the programmers lacked Article III standing because the risk that Copilot would reproduce their specific code, rather than someone else's, was mere conjecture. The panel disagreed. According to the opinion, the complaint cited academic research finding that large language models sometimes emit memorised training data verbatim, and pointed to a GitHub feature that lets users block suggestions matching public code, specifically verbatim snippets of 150 characters or more. The court treated the existence of that duplicate-detection filter as "some evidence that Copilot can and does emit literally identical copies of code." Combined with examples of Copilot reproducing portions of the named plaintiffs' code, that was enough to plead a substantial risk of injury at the motion-to-dismiss stage.
The court was careful to limit the finding. At summary judgment, according to the opinion, the plaintiffs would need evidence creating a genuine factual issue on that risk, and the panel expressly declined to decide whether the evidence referenced in the complaint would meet that standard.
Why the output theory failed
On the merits, the panel agreed with the district court but rejected part of its reasoning. Judge Miller began with dictionary definitions from Webster's Third New International Dictionary: to remove is to get rid of, and to alter is to cause something to become different in some particular characteristic. Both verbs, the court held, imply an affirmative act directed at CMI attached to a work that already exists.
"One who creates a new work and fails to include CMI cannot be said to have 'removed' or 'altered' anything," the opinion states. The statutory definition reinforces the reading, according to the court, because CMI is information conveyed "in connection with copies" of a work, not with excerpts or derivative works.
The plaintiffs' own description of Copilot sank their claim. According to the complaint as quoted by the court, the model infers "statistical patterns governing the structure of code" and predicts the most likely completion to a prompt through a "complex probabilistic process." That account, the panel held, describes a process of generating new works, not an action taken against CMI on an existing file. The opinion draws an explicit contrast with a traditional search engine, which retrieves and displays copies of materials that already exist. Had Copilot worked like a search engine and returned identical copies stripped of notices, "then plaintiffs might have a stronger claim," according to the court. That is not what the plaintiffs alleged.
The identicality question
The certified question asked whether Sections 1202(b)(1) and (b)(3) impose an identicality requirement. The panel's answer is more nuanced than either side's framing. According to the opinion, identicality is "something of a misnomer" because the DMCA does not require literal identity between the original and the allegedly infringing work. The concept is better understood, the court said, as a gloss on the statutory words remove, alter and copies rather than an independent element of the claim.
Identity still has evidentiary weight. Where two works share the same composition, cropping or other distinctive features but one lacks the CMI, a factfinder may infer that the notice was removed. The panel cited its own 2016 decision in Friedman v. Live Nation Merchandise, where photographs were exact copies of images in a book and the only material difference was the missing CMI. It also cited a Third Circuit case in which a reprinted image cropped out a printed gutter credit identifying the photographer, and a Central District of California case in which a photographer who shot a mural from an angle that hid the artist's signature was found not to have removed CMI.
The plaintiffs warned that a strict identicality rule would let infringers escape liability by changing a single word on a single page. The defendants conceded the point, and the court agreed that minor cosmetic changes will not necessarily protect someone who substantially reproduces a work and deletes its notices. The panel quoted a 2024 District of Columbia decision observing that it "would be odd if a defendant could evade DMCA liability by removing or altering CMI in a copied work but only disseminating 99% rather than 100% of that work." It also cited the Southern District of New York's 2025 ruling in New York Times Co. v. Microsoft Corp. on the same point.
The concession did not help the plaintiffs, because they had not alleged removal of CMI from a copy of their work, even an altered one. Sona Sulakian, VP of AI product and co-founder of Pincites, summarised this caveat in a LinkedIn post: "Slightly modifying copied material does not automatically avoid liability."
The damages arithmetic behind the ruling
The final paragraph of the opinion explains why the panel was reluctant to stretch Section 1202(b). Many copyright cases, according to the court, involve a work substantially similar to the plaintiff's that lacks attribution. If that alone triggered the DMCA, the statute "would supplant traditional copyright protections and subject defendants to potentially ruinous liability." The opinion sets the two damages regimes side by side: 17 U.S.C. Section 1203(c)(3) permits up to $25,000 per violation of the DMCA provision, while Section 504(c)(1) caps ordinary statutory damages at $30,000 per work.
The distinction between per violation and per work is the point. A model that emits the same snippet thousands of times could generate thousands of violations under the DMCA reading, but only one infringed work under the Copyright Act. "We decline plaintiffs' invitation to transform run-of-the-mill copyright-infringement claims into DMCA claims," Judge Miller wrote.
How large was the exposure? The opinion does not put a figure on it. According to Sulakian, the plaintiffs sought more than $9 billion in penalties. That figure does not appear in the appellate opinion itself.
The court also flagged the limits of its own analogies. Examples drawn from print media, such as defacing the title page of a book, "do not map neatly onto the emerging digital technologies like artificial intelligence to which the DMCA's protections also apply," according to the opinion. The panel expressed no view on whether Copilot's output could support an ordinary copyright infringement claim.
Who else was in the room
The case drew amicus briefs from both sides of the AI copyright divide. According to the opinion's counsel listing, The Authors Guild, the Association of American Publishers, the News/Media Alliance and the International Association of Scientific Technical and Medical Publishers filed jointly. The Electronic Frontier Foundation and Public Knowledge filed together, as did the Chamber of Progress and the Computer & Communications Industry Association. ACT The App Association, the Authors Alliance and a group of intellectual property law professors also filed.
Jesse Panuccio of Boies Schiller Flexner argued for the programmers, alongside counsel from Butterick Law and the Saveri Law Firm. Lisa S. Blatt of Williams & Connolly and Christopher J. Cariello of Orrick Herrington & Sutcliffe argued for the defendants, with Morrison & Foerster also on the brief.
The Justice Department's brief, 15 days earlier
The Ninth Circuit ruling followed a filing of a different kind in Manhattan. On September 1, 2026, the United States filed a 20-page statement of interest, Document 316, in In re: OpenAI, Inc. Copyright Infringement Litigation, the multidistrict proceeding numbered 25-md-3143 in the Southern District of New York before Judge Sidney H. Stein. The document carries a filing stamp and signature date of September 1. Earlier PPC Land coverage dated the Department's intervention to September 3; the filing itself shows the earlier date.
The statement was submitted under 28 U.S.C. Section 517, which allows the Department of Justice to attend to the interests of the United States in any pending federal suit without the court's leave. It is signed by Associate Attorney General Stanley E. Woodward, Jr., Assistant Attorney General Brett Shumate of the Civil Division and Michael Weisbuch, senior counsel to the Associate Attorney General.
Its conclusion is unambiguous. According to the Department, "the United States has a strong interest in this Court rejecting any argument that training LLMs on copyrighted texts violates copyright law." The brief refers only to The New York Times and OpenAI for simplicity, but states in a footnote that its legal arguments apply to all parties in the litigation and related cases, including book authors and publishers.
National security and the executive orders
The Department grounds its interest in policy documents from the current administration. It cites the President's Executive Order of January 23, 2025, "Removing Barriers to American Leadership in Artificial Intelligence," and a second order of June 2, 2026, "Promoting Advanced Artificial Intelligence Innovation and Security." It also draws on a Government Accountability Office report of April 19, 2022, which warned that failure to integrate AI could hinder national security. According to the brief, rules of law that make it significantly harder to build an AI industry in the United States "threaten national security and give a competitive advantage to foreign adversaries who are not so encumbered."
The brief also cites the March 2026 legislative recommendations in the National Policy Framework for Artificial Intelligence, which, according to the Department, encouraged Congress to consider licensing frameworks or collective rights systems for rights holders while suggesting that such legislation not address when or whether licensing is required.
Training as a separate use
The legal core of the brief is a distinction that the Ninth Circuit would echo two weeks later: training and output are different uses and must be analysed separately. The Department describes LLM development in the three stages set out in an April 4, 2025 opinion in the same court: acquisition, training and output. Training, according to the brief, involves making copies of works and converting them into sets of numbers that let the model learn linguistic patterns. The Department addresses only that middle stage.
On the first fair use factor, the brief calls the copying of text articles for training "extraordinarily transformative" and argues that OpenAI's commercial purpose "does not move the needle." It leans heavily on the June 2025 decisions in Bartz v. Anthropic and Kadrey v. Meta Platforms, both from the Northern District of California. On the second and third factors, a footnote argues that training does not make any copied work accessible to the public, so the copying of entire works at an early step carries little weight.
The Department accepts that outputs are a different matter. According to the brief, certain output uses may not be transformative if an LLM reconstructs and disseminates an original work. Its position is that "whatever legal questions certain output uses might raise, that should not bear on the transformative nature" of the training use. A footnote adds that any remedy would need to be limited to what is necessary for complete relief, and that "[a] tiny sliver of anomalous reconstructive outputs" would not support a remedy threatening LLM output or training generally. The same footnote records that OpenAI states it has taken steps to prevent substantial reproduction of protected material.
The attack on market dilution
The most pointed section concerns the fourth factor, market effect. The Department singles out the Kadrey court's discussion of market dilution, the theory that AI-generated books could flood the market and harm human authors even without copying their expression. According to the brief, the Kadrey court adopted that theory "Without the benefit of briefing," and its application of the fourth factor is "deeply flawed" because it collapsed training and outputs into a single use and applied a genre-level understanding of substitution.
To illustrate, the brief turns to literary history. Joan Didion, as a teenager, typed out Ernest Hemingway's stories to learn how his sentences worked. By the Kadrey logic, according to the Department, Didion would have owed Hemingway every time she published. The brief quotes the Bartz court's view that making anyone pay each time they later draw on a book "when writing new things in new ways would be unthinkable."
The Department also takes aim at the US Copyright Office. A footnote notes that the Register of Copyrights, "who is currently challenging her removal," appeared to endorse a similar dilution theory in the Office's May 2025 report on generative AI training. Citing the Supreme Court's 2024 decision in Loper Bright Enterprises v. Raimondo, the brief states that her understanding "does not warrant deference."
Licensing, legacy media and the independents
The passage most likely to draw attention from publishers concerns who would benefit from mandatory licensing. According to the brief, an erroneous fair use ruling would hamper competition because only the largest technology companies might afford licensing fees, and those fees "would disproportionately benefit legacy media outlets due to the sheer volume of their written publications." The Department describes such barriers as functioning "primarily as large subsidies for old mainstream media companies."
A footnote tempers the claim. The United States "takes no position on whether a licensing regime would be financially or logistically feasible," according to the brief, and it acknowledges that both mainstream and independent publishers have entered licensing agreements to give developers specialised access to real-time, paywalled and proprietary content.
The brief turns The New York Times' own practices against it. Citing a May 12, 2026 report in Futurism, the Department states that Times authors are using LLMs to help them "conceptualize and edit" articles. It also cites an independent writer, a former physics teacher, who used an LLM to publish a contrarian analysis of data centre water use that criticised a July 2025 Times article. The Department argues that such uses show LLMs helping independent publishers compete with mainstream ones.
Where the two documents meet
The two filings arrive from different directions. One is an appellate opinion on a narrow statutory provision, binding on federal courts in the nine western states and two territories of the Ninth Circuit. The other is an advocacy document from a non-party, persuasive at most, filed in a district court in New York. Neither decides the other.
Both, however, rest on the same division between what happens during training and what the model later produces. The Ninth Circuit's refusal to treat generated code as a scrubbed copy tracks the Department's insistence that training and outputs be analysed use by use. And both leave the training question formally open in the matters before them: the Ninth Circuit because the input theory was forfeited, and the Department because a statement of interest cannot bind Judge Stein.
Sulakian put the shift in terms of questions litigants now face: "What happened during training? What did the model actually generate? Was attribution removed, or was it never there?"
The overlap is sharper than a shared philosophy. The New York Times litigation includes its own Section 1202(b)(1) claim, and OpenAI's September 4 summary judgment brief argues that text-extraction scripts shed copyright notices as a side effect of stripping markup from raw web data, not as a target. The Ninth Circuit's reasoning concerns outputs rather than extraction pipelines, and it is not binding in the Second Circuit, where Judge Stein sits. It is nonetheless an appellate reading of Section 1202(b) applied to a generative model, and the parties in New York may cite it.
Why this matters for marketers and publishers
For the advertising and publishing industries, the rulings shift pressure away from one legal tool and onto others. The DMCA route was attractive to rights holders precisely because of the per-violation arithmetic the Ninth Circuit described. With that route narrowed for generated outputs, the contest returns to traditional infringement and fair use, where the evidence is about memorisation and verbatim reproduction.
That fight is already under way in Manhattan. PPC Land reported that OpenAI and Microsoft asked Judge Stein to end the 10.8 million-article case on September 4, citing 24 verbatim outputs in 20 million chat logs. The publishers' experts, by contrast, used adversarial prompting to trigger partial regurgitation of fewer than 2.6 percent of asserted works. Those numbers now carry more weight, because the DMCA shortcut that might have bypassed them is harder to take.
The chat logs themselves exist because a magistrate judge ordered OpenAI to turn over 20 million ChatGPT conversations in November 2025. Publishers also have precedent on the DMCA side that predates the Ninth Circuit: in November 2024 a federal court dismissed Raw Story's DMCA case against OpenAI for lack of standing, in a dispute over CMI allegedly stripped from more than 400,000 articles before training. The Ninth Circuit reached the opposite result on standing for the Copilot plaintiffs, but the same result on the claim.
The Department's brief also lands on a live commercial question. Licensing talks between publishers and AI developers are priced against litigation risk, and the largest public reference point remains Anthropic's $1.5 billion settlement with authors in September 2025, a case in which Judge William Alsup had found training to be fair use but rejected the piracy defence. A federal government arguing publicly that licensing fees amount to subsidies for legacy media does not change the law, but it enters the negotiating record. The New York Times, whose chief executive has said the company has sued three AI companies, has put the annual cost of its journalism at close to $2 billion.
The Copyright Office thread matters too. The Department's dismissal of the Register's views targets the reasoning in the May 2025 report on generative AI training, which had treated market effects as a significant factor. The White House's own national AI policy framework, published in March 2026, recommended that training on copyrighted material remain a question for courts rather than Congress. The Department is now telling a court how to answer it.
For agencies and brands, the Ninth Circuit's CMI reasoning has a practical reach beyond AI. The examples the court cited as paradigmatic violations are ordinary marketing operations: reprinting an image with the photographer's credit cropped out, or deleting metadata that accompanies a digital file. The opinion confirms that those acts remain squarely within Section 1202(b). What falls outside it is the generation of new material that never carried the notice in the first place.
Code is part of this too. Ad tech runs on open-source components, from header bidding wrappers to measurement libraries, much of it under licences whose main condition is attribution. The Ninth Circuit held that a model producing similar code without that attribution does not violate the DMCA. It did not hold that such output is lawful. The contract claims in Doe v. GitHub, which rest on the licence terms themselves, remain pending before Judge Tigar.
What happens next
In the Copilot case, the matter returns to the Northern District of California on the surviving contract claims. The plaintiffs could seek rehearing en banc or petition the Supreme Court, though the opinion gives no indication of either. The panel noted that the Fifth Circuit reached a similar view on the failure to include CMI in Kipp Flores Architects v. AMH Creekside Development, decided on August 21, 2026, which reduces the likelihood of a clean circuit split on this narrow point.
In New York, cross-motions for summary judgment are pending before Judge Stein, and both defendants have requested oral argument. No date has been set, according to earlier PPC Land reporting. The Department's statement will sit alongside those motions. Judge Stein is under no obligation to follow it.
Neither document settles the question that the industry most wants answered. Is copying works to train a model fair use across the board, or does it depend on what the model later does with them? One court declined to reach it on procedure. The executive branch answered it without the power to decide it.
Timeline
- 1998 - Congress enacts the Digital Millennium Copyright Act, including Section 1202 on copyright management information (Pub. L. No. 105-304)
- April 19, 2022 - Government Accountability Office warns that failing to integrate AI could hinder national security
- 2022 - Anonymous programmers file a putative class action against GitHub, Microsoft and OpenAI in the Northern District of California (docket 4:22-cv-06823-JST)
- November 7, 2024 - A federal court dismisses Raw Story's DMCA case against OpenAI for lack of standing
- January 23, 2025 - Executive Order "Removing Barriers to American Leadership in Artificial Intelligence"
- April 4, 2025 - Southern District of New York describes the acquisition, training and output stages of LLM development in New York Times Co. v. Microsoft Corp.
- May 2025 - US Copyright Office publishes Part 3 of its AI report on generative AI training
- June 23, 2025 - Judge Alsup finds AI training fair use in Bartz v. Anthropic but rejects the piracy defence
- June 25, 2025 - Judge Chhabria grants Meta summary judgment on fair use in Kadrey
- September 5, 2025 - Anthropic agrees to a $1.5 billion settlement with authors
- November 13, 2025 - Court denies OpenAI's request to pause production of 20 million ChatGPT logs
- February 11, 2026 - Ninth Circuit hears argument in Doe v. GitHub in San Francisco
- March 2026 - White House publishes its National Policy Framework for Artificial Intelligence
- May 12, 2026 - Futurism reports on The New York Times' warning to freelancers about AI use, later cited by the Department of Justice
- June 2, 2026 - Executive Order "Promoting Advanced Artificial Intelligence Innovation and Security"
- August 2026 - New York Times chief executive says the company has sued three AI companies
- August 21, 2026 - Fifth Circuit decides Kipp Flores Architects v. AMH Creekside Development; The New York Times files the pleading cited in the Department's brief
- September 1, 2026 - Department of Justice files its 20-page statement of interest in 25-md-3143, supporting fair use for LLM training
- September 4, 2026 - OpenAI and Microsoft move for summary judgment in the 10.8 million-article case
- September 16, 2026 - Ninth Circuit affirms dismissal of the DMCA claim against GitHub, Microsoft and OpenAI
Related PPC Land coverage
- OpenAI and Microsoft ask judge to end 10.8 million-article copyright case - The September 4, 2026 summary judgment filings, including OpenAI's answer to the publishers' Section 1202(b)(1) claim.
- OpenAI asks a judge to end the 10.8 million-article copyright case - Analysis of the verbatim output counts at the centre of the defendants' motion.
- Explaining regurgitation - How verbatim reproduction by AI models is measured and why the figures now drive the fair use fight.
- Court dismisses Raw Story's landmark DMCA case against OpenAI over training data - The November 2024 standing decision on the same DMCA provision.
- Court finds fair use in AI training but rejects piracy defense - The Bartz v. Anthropic ruling the Department of Justice relies on.
- Court rules Meta used copyrighted books legally for AI training - The Kadrey decision whose market dilution reasoning the Department attacks.
- Anthropic agrees to $1.5 billion settlement in largest copyright case - The settlement that set the sector's public price benchmark for unlicensed use.
- US Copyright Office releases major AI training report amid intensifying copyright debate - The May 2025 report whose reasoning the Department says deserves no deference.
- White House AI framework targets state laws, child safety, and copyright - The March 2026 legislative recommendations cited in the Department's statement.
- OpenAI must turn over 20 million ChatGPT conversations to New York Times - The discovery order that produced the conversation sample now used in the fair use dispute.
- New York Times has sued three AI companies, CEO Kopit Levien says - The Times' chief executive on litigation and licensing terms with model developers.
- Explaining open source - How licence families and attribution conditions work in ad tech code.
Summary
Who: Anonymous programmers who published code on GitHub, suing GitHub, Microsoft and OpenAI; a Ninth Circuit panel of Judges Sidney R. Thomas, Eric D. Miller and Stanley Blumenfeld, Jr.; and the United States Department of Justice, filing in the New York Times litigation against OpenAI and Microsoft.
What: The Ninth Circuit affirmed dismissal of the programmers' DMCA Section 1202(b) claim, holding that Copilot and Codex generate new works rather than removing copyright management information from copies, while treating the training-stage theory as forfeited. Separately, the Department of Justice filed a 20-page statement arguing that training LLMs on copyrighted text is fair use and criticising the market dilution theory and the Copyright Office's reasoning.
When: The Department of Justice statement is dated and stamped September 1, 2026. The Ninth Circuit opinion was filed on September 16, 2026, after argument on February 11, 2026.
Where: The United States Court of Appeals for the Ninth Circuit, on appeal from the Northern District of California (No. 24-7700), and the United States District Court for the Southern District of New York (No. 25-md-3143).
Why: The DMCA provision carries statutory damages of up to $25,000 per violation, against a $30,000 per-work cap for ordinary infringement, which made it a potent tool against AI tools that emit code or text without attribution. The ruling narrows that route for generated outputs, and the Department's brief adds federal weight to the argument that training is fair use, shifting the contest toward evidence of verbatim reproduction and toward licensing negotiations between publishers and AI developers.
Discussion