Citation: Matthew Sag, The False Hope of Content Licensing at Internet Scale, ProMarket (Nov. 19, 2025)
In a nutshell:
In The False Hope of Content Licensing at Internet Scale, Matthew Sag argues that licensing deals between AI developers and content owners are only possible for large, concentrated rightsholders and cannot feasibly supply the internet-scale training data that cutting-edge large language models require.
Summary
Since mid-2023, OpenAI, Anthropic, Meta, and Google have signed licensing agreements worth hundreds of millions of dollars with the Associated Press, Axel Springer, News Corp, Condé Nast, Hearst, Reddit, and others. Content industry representatives point to these deals as proof that licensing training data is possible, and conclude that AI developers should be required to license. The article explains why these inferences are unsound. Many of the deals are easy to misconstrue: several are more about reliable API access and the delivery of content through retrieval-augmented generation (RAG) than about clearing copyright for model training. Reddit, for example, does not own the copyright in its users’ posts and could never sue AI companies for infringement. Sag acknowledges that agentic uses that quote or closely paraphrase retrieved content at inference time raise different copyright issues than training does, and that there is a plausible case for licensing those uses.
For training itself, however, the article argues that licensing does not scale, for two structural reasons. First, there are not enough dance partners: because the marginal contribution of any individual work to a trillion-token training corpus is approximately zero, licensing is only worthwhile with concentrated media interests holding clear rights to large collections, and the supply of such gatekeepers is limited. Second, the concentrated rightsholders who can overcome these transaction costs do not have enough data. Drawing on earlier work with Peter Yu, Sag notes that it would take the New York Times about 316,000 years to generate the 15 trillion tokens used to train Meta’s Llama 3 model.
The article also rejects collective licensing as a workaround. An ASCAP-style “Large Language Model Rights Association” would fail because current technology provides no reliable way to trace which works influenced a model’s outputs, so compensation cannot be tied to usage. Such an organization would pay the large rightsholders who already have access deals and almost no one else; it would operate more like a tax, or the old system of selling indulgences, than any existing system of collective licensing. The conclusion is careful about what this proves: infeasibility does not settle the fair use question, but it makes clear that cutting-edge LLMs could not be trained on a mixture of public domain and licensed materials alone.
Why read this article?
This short essay is a useful map of the AI training data licensing landscape as of late 2025, cataloging the major deals and explaining what each deal does and does not cover. It offers an accessible explanation of retrieval-augmented generation and why agentic AI raises copyright questions distinct from model training, along with concrete numbers on the scale of modern training corpora. Readers also get a clear account of how collective rights organizations like ASCAP work in the music industry and the specific technical and economic reasons that model does not transfer to LLM training.
Further Reading
Mark A. Lemley & Bryan Casey, Fair Learning, 99 Texas Law Review 743 (2021) – This early and influential article argues that using copyrighted works to train machine learning systems should generally be fair use, in part because requiring permission for training data is impractical and would bias AI systems toward whatever data happens to be cheaply available.
Frank Pasquale & Haochen Sun, Consent and Compensation: Resolving Generative AI’s Copyright Crisis, 110 Virginia Law Review Online 207 (2024) – Pasquale and Sun propose a streamlined opt-out mechanism plus a levy on AI providers to compensate copyright owners, representing the kind of private ordering and collective compensation approach that Sag’s essay argues cannot work at internet scale.
Benjamin L. W. Sobel, Artificial Intelligence’s Fair Use Crisis, 41 Columbia Journal of Law & the Arts 45 (2017) – Written before the generative AI boom, this article anticipated that expressive machine learning applications would strain fair use doctrine and examined the doctrinal and policy options, including licensing, for resolving the tension.