Matthew Sag

DOJ Enters the AI Copyright Fight — and Mostly Gets It Right

So, this just happened. On September 1, the U.S. Department of Justice filed a statement of interest in the sprawling OpenAI copyright litigation, In re OpenAI, Inc., Copyright Infringement Litigation, the multidistrict proceeding before Judge Sidney Stein in the Southern District of New York. It told the court, in essence, that it should hold that training large language models on copyrighted works is fair use.

In its submission, the DOJ explained why training LLMs on copyrighted works is highly transformative and should thus generally be protected by the fair use doctrine. Notably, the DOJ comes to this conclusion by arguing that copyright law should distinguish the training process from any potentially infringing outputs subsequently produced by language models. The statement also contains an important policy overlay, as the Department warns that imposing a broad licensing requirement on AI training would threaten American innovation, competition, and national security.

The DOJ could have done a better job explaining that, although licensing is sometimes an option when we are dealing with millions of rights (think of ASCAP), it simply won’t scale into the billions. As I argued in The False Hope of Content Licensing at Internet Scale, the licensing deals that get the headlines cover a tiny fraction of the data that frontier models actually need. We are seeing the development of a licensing market for access to some specific high-value concentrated resources (like Reddit), but there simply won’t be a licensing market for the trillions of tokens of content needed to train cutting-edge AI models.

I broadly agree with most of what the DOJ put to the court, and much of their argument tracks the non-expressive use principle that I and others have argued for, for well over a decade now. I will come back to where I disagree shortly.

What exactly is a “Statement of Interest”?

The U.S. government is not a party to the litigation involving OpenAI, but filed its statement of interest under 28 U.S.C. § 517, which authorizes the Attorney General to send Justice Department lawyers to “attend to the interests of the United States” in pending litigation. This is basically the government equivalent of an amicus brief. But not just any amicus brief. (Also, unlike a regular amicus brief, a Section 517 statement does not require permission. The government gets to file these as of right.) The federal government typically uses statements of interest only sparingly, and usually where it thinks private litigation implicates important national interests. Another interesting thing about the DOJ’s intervention is that it was signed by the Associate Attorney General and by Senior Counsel to the Associate Attorney General—I don’t know if this is normal, but it doesn’t feel like a routine submission.

The fact of the filing and the position that it takes is a huge victory for OpenAI, and the AI industry more broadly, but the judge is still free to ignore it.

What did the DOJ argue?

I’m going to take their arguments slightly out of order.

(1) AI training is highly transformative

The DOJ’s basic argument is that copying works in the course of training an AI model is highly transformative. Drawing on the summary judgment decisions in Bartz v. Anthropic and Kadrey v. Meta, it wrote:

The copying of protected text articles as part of training an LLM is a use of a different kind or character that is “transformative—spectacularly so.” … An OpenAI LLM thus uses the copyrighted work not to duplicate the work’s expressive content, but as part of a process to learn and act on statistical patterns in written text, including vocabulary, syntax, and knowledge. The purpose of the copying (to build an intelligent, interactive model) differs in kind from the purpose of the copied work (to use language to directly entertain or educate a reading audience). This use of text-based works to create an LLM engine for “innovative tools” that can “edit an email . . . , translate an excerpt from or into a foreign language, write a skit based on a hypothetical scenario, or do any number of other tasks” is undoubtedly “highly transformative.”

The DOJ’s basic argument is hardly radical. It’s basically the same argument that carried the day in cases involving search engines, plagiarism detection, reverse engineering, and computational analysis prior to the generative AI era. Courts have repeatedly found fair use where copyrighted expression must be copied as an intermediate step, but the secondary use does not ordinarily communicate that expression back to the public. The DOJ doesn’t use the term “non-expressive use,” but its argument is exactly in line with the copy-reliant technology cases I have been writing about since 2009 (for a good overview, see Fairness and Fair Use in Generative AI and my forthcoming article, Copyright’s Jagged Frontier).

(2) Market dilution is not a relevant harm

DOJ also persuasively rejects a particularly expansive theory of market harm, “market dilution,” asserted by the Copyright Office and floated by Judge Chhabria in Kadrey v. Meta. It is worth stressing, as the DOJ does, that the Kadrey discussion was dicta offered “without the benefit of briefing” — the court actually granted summary judgment to Meta.

The DOJ’s position is consistent with an amicus brief my co-authors and I filed in Thomson Reuters v. Ross Intelligence, currently before the Third Circuit. Like us, the DOJ argues that copyright does not protect authors against competition as such. If an AI system helps somebody write a different newspaper article, novel, or essay, the mere fact that the new work competes for readers does not make it an infringing substitute. Copyright protects expression, not markets from competition.

In DOJ’s view, market harm becomes relevant to copyright when the secondary use provides a substitute for protected expression, not merely because new technology makes it easier to create competing material. The DOJ’s brief leans on Edward Lee’s Copyright Dilution Under Constitutional Scrutiny for the point that genre-level substitution is not a copyright injury, because “a genre is an uncopyrightable idea or method of expression.”

The DOJ was not holding back in this submission. It made fun of the market dilution argument by pointing out that Joan Didion, as a teenager, typed out Hemingway’s stories to learn how the sentences worked. On the Kadrey dicta’s logic, the government suggests, Didion would have owed Hemingway money every time she published, because the process by which she trained herself and the process by which she produced works were all one use, and her books competed with his in the market for literature. As Bartz put it, requiring payment every time an author “draw[s] upon” a book “when writing new things in new ways would be unthinkable.”

The submission also throws in a pointed footnote, noting that the Register of Copyrights’ endorsement of dilution was based on “threadbare” reasoning and, in any event, warranted no deference under Loper Bright. The fact that the Register is currently challenging her removal did not go unmentioned either.

(3) A narrow fair use ruling would be bad for the United States

According to the DOJ, a restrictive interpretation of the fair use doctrine would not just be bad copyright policy. It would impair national security, weaken American AI companies against foreign competitors, and erect licensing barriers that only the largest technology companies could afford. In short, compulsory licensing would function “primarily as large subsidies for old mainstream media companies.”

This is consistent with the position the administration has taken before, going back to the January 2025 executive order on removing barriers to American leadership in AI and the AI Action Plan. It would be surprising to see a court cite this as the reason for its interpretation of Section 107 of the Copyright Act, but it would not be surprising if courts took account of this below the surface.

(4) Courts should analyze training and output entirely separately

Taking its cue from the Supreme Court’s decision in Andy Warhol Foundation v. Goldsmith, the DOJ argues that fair use analysis “requires an analysis of the specific ‘use’ of a copyrighted work.” In the context of AI training, this means that the copying that took place to assemble the training data and to expose the model to the training data is a separate use from the outputs of the model in generation.

I think this is right, but the DOJ’s argument is a little bit too categorical.

The output and the training copies are legally distinct acts. But evidence about one act can tell us something about the character and consequences of another, at least at the limit.

I agree that evidence that, under abnormal conditions, a model like ChatGPT can be prompted to regurgitate identifiable sections of a New York Times article does not show that training the GPT models was not transformative. Cherry-picked memorization evidence proves much less than its proponents think, as I have argued at length in The Fallacy of Compression. But the separation only holds up to a point. If an AI model was producing infringing output in very large volumes a large percentage of the time in its ordinary operation, I don’t think you could justify the initial training on any kind of transformative fair use theory.

Suppose I train an LLM exclusively on the Harry Potter novels. Then I release a system that, whenever prompted, reliably produces new Harry Potter stories closely reproducing J.K. Rowling’s protected characters, settings, plots, and expression. It would be extraordinarily artificial to say that what this system produces tells us nothing about the nature of the copying that created it.

I doubt that Anthropic or OpenAI would mind conceding this. The whole point of training a large language model is to train on an enormous, diverse corpus; a model that reliably reconstructs its training data is a catastrophically expensive photocopier.

The better framing is the one I develop in Copyright’s Jagged Frontier: training and output are distinct uses, and the fair use analysis should run separately for each, but a developer’s output-side conduct can and should influence the fair use assessment on the training side.

The bottom line

The Department of Justice has told a federal court that training large language models on copyrighted works is transformative fair use, that market dilution is not a cognizable harm, and that a narrow fair use ruling would be bad for the country.

Judge Stein is free to ignore all of this, but hopefully he won’t.


Further reading