Research

My research sits at the intersection of copyright law and technologies that read: search engines, text data mining, machine learning, and generative AI. These pages explain my main contributions in plain English — what each idea is, where it came from, and what has happened because of it. For the whole story in order, see the research agenda; for the works themselves, publications.

Non-Expressive Use (since 2009)

The framework for why machines copying works they never show to humans — search, indexing, mining, AI training — should generally not infringe copyright. Originated in Copyright and Copy-Reliant Technology (2009); now central to the AI training cases.

Copyright and AI Training (since 2023)

Fair use and generative AI, copyright safety for model developers, the “Snoopy problem,” the globalization of training exceptions, and testimony before the U.S. Senate.

Text and Data Mining (since 2012)

Securing researchers’ right to read with machines — from the HathiTrust and Google Books amicus briefs to legal reform proposals in Science.

Fair Use (since 2005)

The doctrine’s structure, its eighteenth-century origins, and the empirical evidence that it is more predictable than its critics claim.

Copyright Trolling (since 2014)

The empirical scholarship that documented and named the mass-litigation business model — cited by federal courts confronting it.

Computational Studies of Supreme Court Oral Argument (since 2007)

With Tonja Jacobi: mining decades of transcripts to measure ideology, interruptions, advocacy, and laughter at the Court.

Impact & Influence

The external evidence: judicial citations, Senate testimony, agency submissions, amicus briefs, citation metrics, and recognition.