Text and data mining (TDM) — also called computational text analysis or corpus analysis — is research that treats large collections of texts or other works as data: counting, comparing, and modeling them rather than reading them one by one. Because TDM usually requires copying the works being analyzed, it raises a copyright question that Matthew Sag has spent more than a decade helping to answer: researchers need the right to read with machines.
The argument
Building on his concept of non-expressive use, Sag has argued since the early 2010s that copying works in order to mine them — for digital humanities scholarship, computational social science, or machine learning research — does not communicate the works’ expression to any human audience and should generally be lawful. He made the case in law reviews (Orphan Works as Grist for the Data Mill, 27 Berkeley Technology Law Journal 1503 (2012)), to scientists in Nature (Digital Archives: Don’t Let Copyright Block Data Mining, 2012, with Matthew Jockers and Jason Schultz), and to the courts in the amicus briefs of digital humanities and law scholars in Authors Guild v. HathiTrust and Authors Guild v. Google — the litigation that established the fair use foundation for library digitization and search in the United States.
The New Legal Landscape for Text Mining and Machine Learning, 66 Journal of the Copyright Society of the U.S.A. 291 (2019), consolidated the post-Google Books doctrine into a practical map for researchers and institutions.
Beyond the United States
Because research is global, Sag has worked with international collaborators to extend TDM rights beyond American fair use: Legal Reform to Enhance Global Text and Data Mining Research, 378 Science 951 (2022) (with Sean Flynn, Jorge Contreras, and others); Implementing User Rights for Research in the Field of Artificial Intelligence: A Call for International Action, European Intellectual Property Review (2020); and submissions to law reform bodies including the Australian Law Reform Commission.
Building the field’s capacity
With support from the National Endowment for the Humanities, Sag was part of the project team for the Building Legal Literacies for Text Data Mining institute at UC Berkeley, and co-authored its open handbook (Building Legal Literacies for Text Data Mining, Pressbooks 2021), which trains librarians and researchers to navigate copyright, contracts, privacy, and ethics in TDM projects. He also served on the HathiTrust Research Center Advisory Board and practices what he preaches: his own computational studies of Supreme Court oral argument are text mining.
Key publications
- Orphan Works as Grist for the Data Mill, 27 Berkeley Technology Law Journal 1503 (2012)
- Digital Archives: Don’t Let Copyright Block Data Mining, 490 Nature 29 (2012)
- The New Legal Landscape for Text Mining and Machine Learning, 66 J. Copyright Soc’y 291 (2019) (SSRN)
- Legal Reform to Enhance Global Text and Data Mining Research, 378 Science 951 (2022) (Science)
- Building Legal Literacies for Text Data Mining (Pressbooks 2021)
Related
Non-expressive use · Copyright and AI training · Computational studies of the Supreme Court · Data sets
