Research Agenda

Last revised August 2026. For the works themselves, see Publications.

The through-line of my research is a single question I started asking before modern AI existed: what should copyright law do about machines that copy works in order to analyze them, rather than to enjoy them? I have pursued that question from the search engines of the 2000s, through text and data mining and machine learning in the 2010s, to generative AI today — alongside a second research program using computational and empirical methods to study courts and copyright litigation. What follows is the story in chronological order.

2005–2009: Fair use theory and the problem of copy-reliant technology

My early work rebuilt fair use from its history and structure upward: God in the Machine (2005) offered a new structural analysis of the fair use doctrine, and The Pre-History of Fair Use (2011) traced the doctrine’s roots back before the famous American cases. This foundation led to the article I am best known for: Copyright and Copy-Reliant Technology (103 Northwestern University Law Review 1607 (2009)), which introduced the concept of non-expressive use. The argument was that technologies which copy expressive works for purposes that do not communicate the work’s expression to any human reader — search engines, plagiarism detection, computational indexing — should generally not be treated as infringing copyright. In 2009 this was an argument about Google; it is now the central contested question in the lawsuits over AI training.

2007–present: The empirical and computational turn

In parallel, with Tonja Jacobi and other co-authors, I helped build the empirical study of intellectual property and of the Supreme Court itself: measuring ideology in IP cases (Ideology and Exceptionalism in Intellectual Property, 97 California Law Review 801 (2009)), and mining the transcripts of Supreme Court oral argument at scale to study interruptions, advocacy by judges, the Chief Justice’s role, and even laughter (The New Oral Argument; Taking Laughter Seriously at the Supreme Court; Supreme Court Interruptions and Interventions). This work is computational text analysis — the same family of techniques whose copyright status my doctrinal work defends. I have practiced what I preach: my scholarship depends on the freedom to mine texts.

The empirical program also transformed how we understand copyright enforcement. Predicting Fair Use (2012) tested what actually drives fair use outcomes. Copyright Trolling, An Empirical Study (2015) and Defense Against the Dark Arts of Copyright Trolling (2018, with Jake Haskell) documented and named the business model of mass “John Doe” copyright litigation; federal courts have cited this work in confronting it. With Pamela Samuelson, I used empirical evidence to measure how eBay changed copyright injunctions (2023).

2012–2019: Text and data mining and machine learning

As mass digitization moved from litigation to infrastructure, I worked to secure the legal foundation for computational research on copyrighted works — what scientists call text and data mining (TDM). I made the case in legal scholarship (Orphan Works as Grist for the Data Mill, 2012), in Nature (Digital Archives: Don’t Let Copyright Block Data Mining, 2012), and in amicus briefs on behalf of digital humanities and law scholars in Authors Guild v. HathiTrust and Authors Guild v. Google — the cases in which the Second Circuit upheld library digitization and Google Books as fair use. The New Legal Landscape for Text Mining and Machine Learning (2019) consolidated the resulting doctrine, and with international collaborators I carried the argument beyond the United States, including in Science (Legal Reform to Enhance Global Text and Data Mining Research, 2022) and in the NEH-funded Building Legal Literacies for Text Data Mining institute.

2023–present: Generative AI

When generative AI arrived, I had been writing about the copyright status of machine learning for over a decade, and the field’s questions arrived on my desk. Copyright Safety for Generative AI (2023) proposed one of the first frameworks for how model developers can reduce the risk of infringing outputs, introducing what I called the “Snoopy problem” — the tendency of models to memorize and reproduce iconic protected characters. Fairness and Fair Use in Generative AI (2024) assessed the fair use claims of AI developers. The Globalization of Copyright Exceptions for AI Training (2025, with Peter K. Yu) mapped the emerging international landscape, and Copyright’s Jagged Frontier (forthcoming in the Duke Law Journal) continues the project. In July 2023 I testified before the U.S. Senate Judiciary Subcommittee on Intellectual Property on copyright and generative AI, and with Pamela Samuelson and Christopher Sprigman I filed comments in the U.S. Copyright Office’s inquiry on AI and copyright.

Generative AI is also a technology the legal profession and legal education must adapt to, and part of my current work addresses that directly: AI Policies for Law Schools (2025), A Student’s Guide to the Law and Policy of AI (2025), my AI Policy Primer, and my casebook Copyright Law in the Age of AI, which integrates AI throughout the teaching of United States copyright law.

Why it holds together

Seen whole, the agenda is one project: understanding how copyright law should respond when reading becomes something machines do. I made the doctrinal argument for non-expressive use before the technologies that now depend on it existed at scale; I built and used the computational methods the argument protects; I documented empirically how copyright litigation actually behaves; and I now work on the generative AI questions those threads converge on — while helping the legal profession and legal education adapt to the same technology. The convening side of my work — the Legal Scholars Roundtable on Artificial Intelligence and the Copyright Scholars Roundtable — keeps that conversation moving in the field at large.