Matthew Sag

Copyright’s Jagged Frontier

Citation: Matthew Sag, Copyright’s Jagged Frontier, 76 Duke Law Journal (forthcoming 2026)

In a nutshell:

In Copyright’s Jagged Frontier, Matthew Sag argues that the boundary between infringing and non-infringing uses of generative AI will be jagged rather than smooth, and that AI companies that adopt “copyright safety” measures to monitor and filter their models’ outputs will both preserve their fair use defenses and lay the foundation for a licensing industry on the model of YouTube’s Content ID.

Summary

The article applies Ronald Coase’s insight that legal rules are starting points for adaptation and negotiation to the copyright disputes over generative AI. On rights, Sag accepts that AI training has a strong claim to fair use under the nonexpressive use cases and the recent summary judgment rulings in Bartz v. Anthropic and Kadrey v. Meta, but he shows that the claim will falter in particular contexts. Memorization is uneven: the Cooper study found that Llama 3.1 70B memorized nearly all of Harry Potter and the Sorcerer’s Stone, even though most models memorize little from most books. Copyright’s substantial similarity threshold is also uneven. A few overlapping notes can infringe a song, and a recognizable sketch can infringe a copyrightable character (the “Snoopy problem”), while literary works require much closer copying. These two sources of variation compound, so an AI developer could apply nearly the same training process to a music model and a text model and face very different legal exposure.

On adaptation, the article argues that AI companies are more exposed than they assume. The courts in the AI training cases treated safeguards against infringing outputs as material to fair use, so fair use depends in part on what a system produces in practice, and too many infringing outputs would undermine the transformative use argument. Sag also questions the industry assumption that the volitional conduct doctrine shields developers from direct liability for infringing outputs. A close reading of the case law, from Netcom through CoStar v. LoopNet, Fox v. Dish, and VHT v. Zillow, shows that the doctrine functions as a proximate cause inquiry turning on foreseeability and reasonable precautions. Prompt filtering and output filtering are therefore a practical necessity for managing fair use, direct liability, and secondary liability risk.

The negotiation section makes the article’s most counterintuitive claim. Filtering will overblock lawful expression and threatens the democratizing potential of generative AI. Yet the history of YouTube’s Content ID (over $12 billion in cumulative payouts to rights holders, with more than 90 percent of claims monetized rather than blocked) shows that copyright safety infrastructure tends to become the foundation of licensing markets. Sag reads the Suno and Udio settlements with the major labels, the Disney-OpenAI deal, and Google’s Lyria 3 release as early evidence that this pattern is repeating. The conclusion generalizes the point: copyright is reframing the generative AI question as one of managed risk within “zones of contingent permission,” a template likely to recur in other domains of AI governance.

Why read this article?

Copyright’s Jagged Frontier brings together material that is otherwise scattered across separate doctrinal and technical literatures. The article provides a plain-English account of the computer science research on memorization, including the Carlini and Cooper studies and the tradeoff between memorization and learning. It surveys substantial similarity doctrine, showing how loosely music cases like Bright Tunes and Blurred Lines apply the test compared to literary cases like Williams v. Crichton, and it gives a sustained treatment of the volitional conduct doctrine in the generative AI context, from Netcom through Aereo to VHT v. Zillow. The article also traces the history of the DMCA safe harbors, DMCA-plus agreements, and Content ID, and analyzes the AI licensing deals announced in late 2025 and early 2026, several of which closed as the article was being written.

Further Reading

Mark A. Lemley & Bryan Casey, Fair Learning, 99 Texas Law Review 743 (2021) – An early and influential argument that training machine learning systems on copyrighted works should generally be fair use because the systems copy works for their ideas and functional content rather than their expression.

A. Feder Cooper et al., Extracting Memorized Pieces of (Copyrighted) Books from Open-Weight Language Models (2025) – The empirical study at the center of the article’s memorization discussion, showing that while most open-weight models memorize little from most books, some models memorize some books almost entirely.

Nicholas Carlini et al., Quantifying Memorization Across Neural Language Models, ICLR 2023 – Demonstrates that memorization in language models grows with model size, with the number of times an example is duplicated in the training data, and with the length of the prompting context.

Katrina Geddes, Engineering Semiotic Democracy, FIU Law Review (forthcoming 2026) – A critical account of how copyright filters on generative AI platforms suppress lawful user expression, including many fair uses; the article engages directly with Geddes’s argument in assessing the costs of filtering.