Citation: Matthew L. Jockers, Matthew Sag & Jason Schultz, Digital Archives: Don’t Let Copyright Block Data Mining, 490 Nature 29 (2012)
In a nutshell:
In Digital Archives: Don’t Let Copyright Block Data Mining, Matthew L. Jockers, Matthew Sag and Jason Schultz argue that copying books for nonexpressive purposes such as text mining and computational analysis does not infringe copyright, and that the courts in the Google Books litigation should say so explicitly.
Summary
This short comment, written for Nature’s scientific audience at the height of the Google Books litigation, explains why a group of 64 scholars from law, computer science, linguistics, history and literature (including the three of us) filed amicus curiae briefs in Authors Guild v. Google and Authors Guild v. HathiTrust. The article recounts the case history: Google began scanning library collections in 2004 and had digitized more than 20 million books, most of them out of print; the Authors Guild sued in 2005; a proposed class-action settlement collapsed in 2011; and a parallel suit targeted the universities behind the HathiTrust Digital Library. If the Authors Guild prevailed, digital humanities scholars might have been limited to analyzing public-domain works, which would cut off as much as two-thirds of the literary record.
The legal argument is an early, compact statement of what Sag has elsewhere developed as the theory of nonexpressive use. Copyright protects an author’s original expression, not the facts and ideas contained in it. Text mining converts masses of text into metadata such as word frequencies and search indexes, and no human reader ever experiences the author’s expression. Unauthorized music file sharing can infringe because people ultimately experience the files as musical works, but scanning library books to build a search index or count words does not interfere with any interest copyright protects. Courts had already accepted this logic for web search engines and for Turnitin’s plagiarism-detection copying.
The article also shows what was at stake for research. Clustering more than 3,000 nineteenth-century novels by stylistic and thematic features reveals that books by men cluster separately from books by women, with George Eliot’s novels sitting firmly among the male authors. Findings like that cannot be gleaned from close reading a handful of books. We conclude by urging the courts to hold that digitization for text mining and computational analysis is fair use.
Why read this article?
At two pages, this is a quick introduction to the copyright issues that surrounded the Google Books cases, and a historical document in its own right: it was published shortly before the HathiTrust and Google Books fair use decisions adopted much of the reasoning it advocates. The article gives a short timeline of the litigation and the failed settlement, and explains the idea-expression distinction in plain language for scientists rather than lawyers. It also provides a concrete example of macroanalysis in the digital humanities, including a network visualization of 3,000 novels. The same nonexpressive use argument now anchors the debate over copyright and AI training data, so the article doubles as a record of where that argument started.
Further Reading
Ian Hargreaves, Digital Opportunity: A Review of Intellectual Property and Growth (UK Intellectual Property Office, 2011) – This independent review commissioned by the British government recommended copyright exceptions for text and data mining, reaching conclusions parallel to the article’s on the economic case for allowing computational analysis.
Pamela Samuelson, The Google Book Settlement as Copyright Reform, 2011 Wisconsin Law Review 479 – Samuelson analyzes how the proposed Google Books settlement would have accomplished copyright reforms that Congress would find difficult to enact, and asks whether that quasi-legislative character counted for or against judicial approval.
James Grimmelmann, The Elephantine Google Books Settlement, 58 Journal of the Copyright Society of the U.S.A. 497 (2011) – Grimmelmann examines the class action, copyright and antitrust dimensions of the settlement, arguing that its significance lay in fusing these legal categories into a novel mechanism for restructuring an intellectual property industry.
Franco Moretti, Graphs, Maps, Trees: Abstract Models for a Literary History (Verso, 2005) – Moretti’s manifesto for “distant reading” supplies the intellectual background for the macroanalytic methods the article defends, charting whole genres and national literatures rather than close reading individual texts.