Non-Expressive Use

Non-expressive use is the principle that copying a work for purposes that do not communicate its expression to any human reader — indexing, search, plagiarism detection, text and data mining, machine learning — should generally not infringe copyright. Matthew Sag originated the concept in Copyright and Copy-Reliant Technology, 103 Northwestern University Law Review 1607 (2009), and it now sits at the center of the copyright disputes over AI training.

The idea in plain English

Copyright exists to protect an author’s expression as it is experienced by people. But modern technologies copy works for reasons that have nothing to do with human experience: a search engine copies the web to index it; a plagiarism detector copies essays to compare them; a language model ingests text to learn statistical patterns from it. When no human ever reads the copy, the copying does not communicate the author’s original expression to a new audience — and so, the argument runs, it does not implicate copyright’s core concern. In ordinary language: machines that read for us, rather than to us, are doing something copyright law should generally permit.

How the idea developed

  • 2009Copyright and Copy-Reliant Technology (Northwestern University Law Review) introduces the framework, arguing that copy-reliant technologies such as search engines make non-expressive use of copyrighted works and should generally prevail on fair use.
  • 2012 — The argument extends to mass digitization and the digital humanities in Orphan Works as Grist for the Data Mill (Berkeley Technology Law Journal) and, in Nature, “Digital Archives: Don’t Let Copyright Block Data Mining” (with Matthew Jockers and Jason Schultz).
  • 2012–2014 — Amicus briefs of digital humanities and law scholars, authored with colleagues, put the framework before the courts in Authors Guild v. HathiTrust and Authors Guild v. Google — the cases in which the Second Circuit upheld library digitization and Google Book Search as fair use.
  • 2019The New Legal Landscape for Text Mining and Machine Learning (Journal of the Copyright Society) maps the resulting doctrine for the machine learning era.
  • 2023 — Sag testifies before the U.S. Senate Judiciary Subcommittee on Intellectual Property on copyright and generative AI; Copyright Safety for Generative AI (Houston Law Review) addresses how model developers can stay on the right side of the line.
  • 2025–2026 — The framework is contested terrain in the generative AI litigation wave, engaged in The Globalization of Copyright Exceptions for AI Training (Emory Law Journal, with Peter K. Yu), an amicus brief of copyright law professors in Thomson Reuters v. Ross Intelligence (3d Cir.), and Copyright’s Jagged Frontier (Duke Law Journal, forthcoming).

Key publications

  • Copyright and Copy-Reliant Technology, 103 Northwestern University Law Review 1607 (2009)
  • The New Legal Landscape for Text Mining and Machine Learning, 66 Journal of the Copyright Society of the U.S.A. 291 (2019) (SSRN)
  • Copyright Safety for Generative AI, 61 Houston Law Review 295 (2023) (SSRN)
  • Fairness and Fair Use in Generative AI, 92 Fordham Law Review 1887 (2024) (SSRN)

Evidence of influence

  • Invited testimony before the U.S. Senate Judiciary Subcommittee on Intellectual Property on copyright and generative AI (July 2023).
  • Amicus briefs in Authors Guild v. HathiTrust, Authors Guild v. Google, and Thomson Reuters v. Ross Intelligence.
  • Comments and reply comments (with Pamela Samuelson and Christopher Jon Sprigman) in the U.S. Copyright Office’s Notice of Inquiry on Artificial Intelligence and Copyright (2023).
  • The concept has been adopted, debated, and extended in subsequent copyright scholarship on search engines, text and data mining, and AI training.

Related

Text and data mining · Copyright and AI training · Fair use · Research agenda