The United States Department of Justice has entered the copyright war over artificial intelligence on the side of the model makers. In a Statement of Interest in the consolidated In re OpenAI Inc. Copyright Infringement Litigation in the Southern District of New York, the government told Judge Sidney H. Stein that training a large language model on copyrighted news articles is a transformative fair use that should not expose developers to broad infringement liability. Signed by senior Justice Department official Stanley Woodward Jr., the 20-page brief landed at the start of September, days before The New York Times and rival publishers moved to reject that defense.
Reuters reported on September 2 that the filing appeared to be the first time the federal government had weighed in on the wave of AI copyright cases brought by authors, publishers, music labels, and news outlets. The intervention is advisory, not binding. A Statement of Interest, rooted in 28 U.S.C. Section 517, lets the Attorney General alert a court to the interests of the United States in a case to which the government is not a party. The signal is unambiguous: the administration regards training AI models on copyrighted text as lawful copying and frames a ruling to the contrary as a threat to American competitiveness and national security.
Key Facts
The case is a multidistrict consolidation, No. 25-md-3143, gathering claims by The New York Times, the New York Daily News, the Chicago Tribune, other newspaper owners, and authors including George R.R. Martin and John Grisham against OpenAI and its largest backer, Microsoft. The Times filed suit in December 2023, accusing the companies of copying millions of its articles to build ChatGPT. Lowenstein Sandler reported on September 4 that the department urges the court to separate how training data is acquired, how a model is trained, and what it generates, assessing fair use stage by stage rather than against the system as a whole.
The core argument is that internal copying during training is transformative to an extraordinary degree; coverage calls the government's characterization exceedingly and even spectacularly transformative. Law360 reported on September 2 that the Justice Department was telling the judge that using copyrighted texts to train generative tools should not expose developers to broad infringement liability. A model does not reproduce an article for a reader, the government says; it reduces text to statistical patterns of vocabulary and syntax, a purpose different in kind from reporting meant to inform or entertain. Commerciality should not be decisive when the use is transformative and exposes no protected expression to the public, the brief adds.
On market harm, the fourth factor, the government argues training copies do not substitute for the originals; relevant harm must come from outputs that displace them. That leads it to attack the market dilution theory credited in Kadrey v. Meta Platforms, under which AI output could flood a genre and depress what authors are paid; the department calls that theory profoundly flawed. The Hill reported on September 2 that administration lawyers wrote that it would be legally incorrect to impose broad copyright liability that would make AI training impermissible without licensing.
Two days later, on September 4, The New York Times, the Daily News, and affiliated papers moved for partial summary judgment, asking Judge Stein to rule that fair use cannot excuse copying at any stage, from paywall circumvention to training to outputs that reproduce protected text. OpenAI and Microsoft filed cross-motions the same day. WebProNews reported on September 5 that Microsoft cited 8.2 million Copilot chat logs, asserting that fewer than 1 percent contained 16 or more words matching published news content.
Analysis
Fair use turns on four factors: the purpose and character of the use, the nature of the work, the amount taken, and the effect on the market. The DOJ tries to win on the first and fourth while avoiding the second. News is factual and close to the core of what copyright encourages, which is awkward for the defendants. But the government insists the copying serves a different purpose, machine learning rather than reader consumption, and that no reader ever receives an article from the corpus.
The brief also exposes a split inside the federal government. The Copyright Office, in its May 2025 report on generative AI training under Register Shira Perlmutter, took a cautious, fact-specific view, warning that commercial training producing competing outputs might fall outside fair use. The executive branch now urges the opposite. Lowenstein Sandler reported on September 4 that the DOJ's stance diverges from the Copyright Office's measured analysis and reflects an administration-level priority, reinforced by the White House National AI Legislative Framework of March 2026, which declared that such training does not inherently violate copyright law and urged Congress to leave the question to the courts.
The bigger picture here is that the executive branch has asked one district judge to settle, through the four fair use factors, a policy dispute Congress has declined to resolve. The national security framing is the most consequential part of the brief and the most fragile as doctrine. Courts weigh the statutory factors; they do not measure the balance of power between Washington and Beijing. A judge who reads the factors for the publishers will not be moved by warnings that a loss would aid Chinese AI rivals, however forcefully the administration asserts them.
What this really means is that the Anthropic settlement and the DOJ brief pull in opposite directions. In July 2026, a California federal judge approved Anthropic's $1.5 billion settlement with authors and publishers over pirated copies of more than 482,000 books used to train the Claude chatbot, a payout of roughly $3,100 per work. The deal shows AI companies will pay enormous sums when they acquire data badly. The DOJ brief argues the opposite: lawfully obtained text can be used to train models without paying its owners. If Judge Stein agrees, licensing becomes a business choice, not a legal requirement. If he rejects that view, model makers owe rights holders for the corpora on which their systems are built.
Why It Matters
The Times and its co-plaintiffs are not litigating only their own archives. A ruling that training on copyrighted journalism is not fair use would reshape the economics of every lab that depends on web-scale text, because the same logic would carry to books and essays across dozens of parallel cases. A ruling for OpenAI and Microsoft would bless the ingestion of copyrighted material as a costless input, leaving creators to rely on output-based claims and on deals like those The Times has already struck with some AI partners.
Graham James, a Times spokesman, criticized the administration for siding with a handful of trillion-dollar AI companies at the expense of the countless American creators whose work, in the publishers' telling, was taken without permission. Steven Lieberman, the publishers' lawyer, says the position ignores the Copyright Clause and contradicts the Copyright Office. The dispute has become a test of whether copyright exists to reward human authorship or to clear a path for machine intelligence.
There is also an enforcement subplot. Publishers have a pending sanctions motion against OpenAI over allegations that it destroyed evidence and concealed its ability to locate news stories in training data and chat outputs. Judge Stein could rule on that motion before the merits, and unsealing of disputed exhibits could reshape the public record either way.
Next Up
Judge Stein must now decide whether any fair use question can be resolved as a matter of law or must go to a jury. Briefing on the dueling motions is underway. If the publishers win in whole or in part, the case proceeds toward trial on damages and on outputs The Times says reproduce its journalism. If OpenAI and Microsoft prevail, the consolidated action could narrow quickly, erasing billions in potential exposure.
The government's brief does not bind him, and district judges have already split on AI training. A federal court in 2025 rejected a fair use defense in Thomson Reuters' case against the legal research company Ross Intelligence, while Judge William Alsup in the Anthropic case suggested that training on lawfully acquired books could be fair use even as he faulted the company for taking books from pirate libraries. No appellate court has ruled definitively, so the loser will have strong reasons to appeal.
The administration's logic also collides with other parts of the copyright landscape. It does not depend on how training data was acquired, which sidesteps the piracy problem behind the Anthropic settlement but leaves unresolved what happens when a model output recites a copyrighted passage verbatim. Europe has no fair use doctrine; the EU AI Act imposes copyright obligations on providers regardless of where training occurred, so a ruling favoring American labs would not export cleanly. For now, the fastest moving forum in the global debate over AI and copyright is a single courtroom in lower Manhattan, and the federal government has told the judge which side it believes should win.
Comments (0)
Log in or sign up to leave a comment.
No comments yet. Be the first to share your thoughts.