Matthew Sag

The New Legal Landscape for Text Mining and Machine Learning

Citation: Matthew Sag, The New Legal Landscape for Text Mining and Machine Learning, 66 Journal of the Copyright Society of the U.S.A. 291 (2019)

In a nutshell:

In The New Legal Landscape for Text Mining and Machine Learning, Matthew Sag argues that the Authors Guild cases settled the fair use status of text data mining in the United States, and maps the legal questions that remain for researchers under contract law, anti-circumvention rules, and foreign copyright law.

Summary

The New Legal Landscape for Text Mining and Machine Learning takes stock of the legal environment for text data mining (TDM) in the United States after the Second Circuit’s decisions in Authors Guild v. HathiTrust and Authors Guild v. Google. Those cases held that digitizing millions of library books to enable full-text search and text mining was transformative and ultimately fair use. Sag explains why that result follows from copyright’s deep structure: the exclusive rights of the copyright owner are defined by, and limited to, the communication of original expression to the public. Non-expressive uses, acts of copying undertaken to derive metadata about works rather than to enable human enjoyment of the expression within them, do not communicate that expression to any human audience and so pose no threat of expressive substitution.

The article traces this principle through earlier caselaw on software reverse engineering (Sega v. Accolade, Sony v. Connectix), image search (Kelly v. Arriba Soft, Perfect 10 v. Amazon.com), and plagiarism detection (iParadigms), and argues that the Authors Guild precedents are likely to remain settled law. It also works through the Second Circuit’s muddled decision in Fox News v. TVEyes, concluding that the case leaves the fair use status of TDM undisturbed while showing that courts will continue to police the amount of original expression displayed to end users.

Parts II and III go beyond the core fair use holding. Part II surveys the treatment of TDM in other jurisdictions, including the statutory exceptions adopted in the United Kingdom, France, Germany, and Estonia and the two mandatory TDM exceptions in the EU’s 2019 Digital Single Market Directive. It also explains how the European Court of Justice’s low threshold for reproduction in Infopaq complicates snippet display, and even some machine learning applications, in Europe. Part III introduces a four-stage model of TDM research (Access, Extraction, Mining, and Use) and applies it to the issues the Authors Guild cases did not resolve, including license terms, the Computer Fraud and Abuse Act, the DMCA’s anti-circumvention provisions, security requirements, and the display of search results and verification snippets.

Why read this article?

The article is written to be useful to lawyers and researchers alike. It opens with examples of what text mining does in practice: literature-based knowledge discovery in the life sciences, Underwood, Bamman, and Lee’s study of gender in over 100,000 novels in the HathiTrust collection, machine learning systems trained on photographs and social media, and the web crawling behind every search engine. It then supplies the doctrinal history behind the non-expressive use argument, from Baker v. Selden and Feist through the Authors Guild litigation, including extracts from the oral argument in the Google Books case.

The comparative material in Part II remains a helpful orientation to the EU’s Digital Single Market Directive and the differences between fair use and enumerated TDM exceptions. The four-stage model in Part III works as a practical checklist for anyone designing a text mining or machine learning project, identifying the legal questions that arise at each stage from acquiring the corpus to publishing the results.

Further Reading

James Grimmelmann, Copyright for Literate Robots, 101 Iowa Law Review 657 (2016) – This essay observes that copyright law has effectively concluded that reading by computers does not count as infringement, and explores the consequences of a system that treats human and robotic readers so differently.

Benjamin L. W. Sobel, Artificial Intelligence’s Fair Use Crisis, 41 Columbia Journal of Law & the Arts 45 (2017) – Sobel examines whether fair use can accommodate expressive machine learning applications, and argues that current doctrine threatens either to impede machine learning or to disadvantage the human creators whose works supply the training data.

Michael W. Carroll, Copyright and the Progress of Science: Why Text and Data Mining Is Lawful, 53 UC Davis Law Review 893 (2019) – Carroll argues that United States copyright law permits researchers to conduct computational analysis of any materials to which they have lawful access, and analyzes how security precautions should figure in the fair use analysis.

Christophe Geiger, Giancarlo Frosio & Oleksandr Bulayenko, Text and Data Mining in the Proposed Copyright Reform: Making the EU Ready for an Age of Big Data?, 49 IIC – International Review of Intellectual Property and Competition Law 814 (2018) – This article analyzes the text and data mining exceptions proposed for the EU’s Digital Single Market Directive and argues for a broader, more innovation-friendly approach to TDM in European copyright law.