Citation: Matthew Sag, Copyright and Copy-Reliant Technology, 103 Northwestern University Law Review 1607 (2009)
In a nutshell:
In Copyright and Copy-Reliant Technology, Matthew Sag argues that acts of copying which do not communicate the author’s original expression to the public should not generally constitute copyright infringement, a principle he names “nonexpressive use.”
Summary
Copyright and Copy-Reliant Technology is the article that originated the concept of nonexpressive use. Sag identifies a class of “copy-reliant technologies” that copy expressive works automatically and in bulk in order to process them as raw material for algorithms and indices. The article develops the concept through four case studies: the search engine caching at issue in Field v. Google, the image-search thumbnails in Perfect 10 v. Amazon, the mass digitization of library books in the Google Book project, and the Turnitin plagiarism detection service. Each involves copying entire works without conveying their expression to a human audience.
The doctrinal argument proceeds from a survey of existing copyright law. The idea-expression distinction, the test of substantial similarity, the collective work right construed in New York Times v. Tasini, and the courts’ refusal to base infringement claims on unpublished drafts all point the same way: the copyright owner’s exclusive rights are implicitly defined and limited by reference to the communication of original expression to the public. Copyright protects authors against expressive substitution, and a use incapable of substituting for the author’s expression falls outside the core of the copyright entitlement. Sag argues that courts should implement the principle through the fair use doctrine rather than as a freestanding defense: nonexpressive uses are equivalent to highly transformative uses under the first factor, qualitatively insignificant under the third, and produce no cognizable market effect under the fourth. Computer software, whose ordinary use is functional rather than expressive, is treated as an exception.
The final Part addresses transaction costs. Copy-reliant technologies need close to complete coverage to be useful, so clearing rights work by work is prohibitively expensive; the article estimates that clearing rights for the Google Book corpus would have cost between one and ten billion dollars before any royalties were paid. Neither digital rights management nor collective rights management solves the problem at Internet scale. Sag concludes that low-cost opt-out mechanisms such as the Robots Exclusion Protocol preserve the autonomy of copyright owners, and that a defendant’s provision of an effective opt-out should weigh in its favor under the first and fourth fair use factors.
Why read this article?
Anyone tracing the intellectual history of the copyright and AI training debate should start here. Published in 2009, well before the machine learning cases, the article supplies the analytical framework that courts later drew on in the HathiTrust and Google Books litigation and that now structures arguments over generative AI. The four case studies also serve as a detailed factual record of how search engines, caching, book digitization, and plagiarism detection actually worked in the 2000s.
The article provides useful doctrinal background as well, working through Baker v. Selden, Feist, Tasini, Campbell, and the software reverse engineering cases to show how expressive substitution organizes seemingly unrelated corners of copyright law. The transaction-cost analysis in Part III anticipates debates that later surfaced in the Google Books settlement objections and the EU’s text and data mining opt-out.
Further Reading
Wendy J. Gordon, Fair Use as Market Failure: A Structural and Economic Analysis of the Betamax Case and Its Predecessors, 82 Columbia Law Review 1600 (1982) – The foundational economic account of fair use as a response to market failure; Part III of Sag’s article builds on this tradition.
Pamela Samuelson, Unbundling Fair Uses, 77 Fordham Law Review 2537 (2009) – Published the same year, this article argues that fair use decisions fall into predictable policy-relevant clusters, including one covering technology and information-access uses.
James Grimmelmann, Copyright for Literate Robots, 101 Iowa Law Review 657 (2016) – A direct engagement with the nonexpressive use idea, arguing that copyright law has consistently treated reading by machines as categorically different from reading by humans.
Maurizio Borghi & Stavroula Karapapa, Copyright and Mass Digitization (Oxford University Press 2013) – A cross-jurisdictional and more skeptical treatment of mass digitization, examining rights clearance, orphan works, and whether copyright can accommodate copying at this scale.