Citation: Matthew Sag, Orphan Works as Grist for the Data Mill, 27 Berkeley Technology Law Journal 1503 (2012)
In a nutshell:
In Orphan Works as Grist for the Data Mill, Matthew Sag argues that there is no orphan works problem for library digitization directed at search and data analysis, because copying expressive works for non-expressive purposes does not infringe the rights of copyright owners.
Summary
Written for the 2012 Berkeley symposium on orphan works and mass digitization, the article separates library digitization into three projects: preservation, distribution, and search. Digitizing books in order to display or distribute them implicates the copyright owner’s exclusive rights, and the proposed settlement in Authors Guild v. Google, which would have let Google sell access to entire books, went well beyond anything fair use could justify. Digitization for search and data processing is different. Because the exclusive rights of copyright owners are limited to the expressive elements and expressive uses of their works, scanning millions of books to build a search engine or a research corpus requires no permission at all, and the impossibility of clearing rights in orphan works is beside the point.
The core of the article is the concept of “non-expressive use”: reproduction as part of a process of automated data analysis that never communicates the author’s original expression to a human reader. Word frequency tables, search indexes, plagiarism fingerprints, and Google Ngram charts are data about books rather than replacements for them. Sag shows that courts already apply this principle without naming it. Substantial similarity is judged from the perspective of the consuming public, and courts have declined to find infringement in intermediate drafts of screenplays that were never shown to an audience. The reverse engineering cases go further and treat wholesale intermediate copying of computer software as fair use. Each doctrine reflects the same premise: copyright protects authors against expressive substitution, and a use that poses no threat of expressive substitution falls outside the author’s entitlement.
The article then argues that the principle should be operationalized through fair use rather than as a freestanding defense, both because the Copyright Act’s definition of reproduction contains no requirement that anyone actually perceive the copy, and because a per se rule would unsettle protection for functional subject matter like computer software. Under the four-factor analysis, non-expressive uses should be treated as equivalent to highly transformative uses on the first factor, qualitatively insignificant on the third, and free of cognizable market harm on the fourth.
Why read this article?
Orphan Works as Grist for the Data Mill is an early statement of the non-expressive use theory that later shaped the fair use holdings in Authors Guild v. HathiTrust and Authors Guild v. Google, and that now anchors debates over copyright and AI training data. The article maps the Google Books litigation as it stood in 2012 and traces the relevant doctrine from Baker v. Selden and Feist through the Seinfeld quiz book and Harry Potter Lexicon cases. It also explains what text mining makes possible, from computational linguistics and automated translation to the macro-analysis of literature practiced by digital humanities scholars such as Franco Moretti and Matthew Jockers, and it collects data on the scale of the orphan works problem, including estimates that most books in U.S. library collections are out of print and earning nothing for their rights holders.
Further Reading
Jennifer M. Urban, How Fair Use Can Help Solve the Orphan Works Problem, 27 Berkeley Technology Law Journal 1379 (2012) – Published in the same symposium issue, this article argues that fair use already permits nonprofit libraries and archives to make some uses of orphan works without waiting for legislation, addressing the expressive-use side of the problem that Sag’s article sets aside.
Pamela Samuelson, The Google Book Settlement as Copyright Reform, 2011 Wisconsin Law Review 479 (2011) – This article shows how the proposed Google Books settlement would have accomplished a form of copyright reform, including a de facto solution to the orphan works problem, that Congress had been unable to deliver.
James Grimmelmann, The Elephantine Google Books Settlement, 58 Journal of the Copyright Society of the U.S.A. 497 (2011) – Grimmelmann argues that the class action, copyright, and antitrust objections to the settlement were all facets of a single underlying concern about concentrating control over digitized books, and that the settlement had to be evaluated as a whole.
Matthew L. Jockers, Macroanalysis: Digital Methods and Literary History (University of Illinois Press, 2013) – This book demonstrates the scholarship that non-expressive use makes possible, using computational analysis of thousands of novels to study literary style, theme, and influence at a scale no close reader could match.