My research sits at the intersection of copyright law and technologies that read: search engines, text data mining, machine learning, and generative AI. These pages explain my main contributions in plain English — what each idea is, where it came from, and how it has developed since. For the whole story in order, see the research agenda; for the works themselves, publications.
Non-Expressive Use (since 2009)
A framework for why machine copying for search, indexing, mining, and AI training should generally not infringe when the copied expression is not communicated to humans. Originated in Copyright and Copy-Reliant Technology (2009); now central to the AI training cases.
Copyright and AI Training (since 2023)
Fair use and generative AI, copyright safety for model developers, the “Snoopy problem,” the globalization of training exceptions, and testimony before the U.S. Senate.
Text and Data Mining (since 2012)
Securing researchers’ right to read with machines — from the HathiTrust and Google Books amicus briefs to legal reform proposals in Science.
Fair Use (since 2005)
The doctrine’s structure, its eighteenth-century origins, and the empirical evidence that it is more predictable than its critics claim.
Copyright Trolling (since 2014)
Empirical research documenting and naming the mass-litigation business model, which federal courts have since cited.
Computational Studies of Supreme Court Oral Argument (since 2007)
With Tonja Jacobi: mining decades of transcripts to measure ideology, interruptions, advocacy, and laughter at the Court.
Impact & Influence
The external evidence: judicial citations, Senate testimony, agency submissions, amicus briefs, citation metrics, and recognition.