I am the co-convenor of two academic roundtables: one devoted to Law and AI, another devoted to copyright. Between them, I read about 60–80 submissions last year that were supposed to be, more or less, complete papers. Only one struck me as pure AI slop. Many others contained what I took to be signs of AI drafting, and I mean that in a bad way.
Of course, I don’t know how any of these individual cakes were baked, but it used to be easier to spot a low quality paper or a paper by someone who hadn’t read the literature they claimed to be contributing to. As competent-sounding prose becomes easier to produce, the contribution of any new paper becomes harder to evaluate.
I was reminded of this over the Labor Day weekend, when “AI slop” was prominent on my X feed and, largely by coincidence, I read two working papers approaching it from very different directions. Both have useful things to say, but neither one really gets to the question of how much work, and what kind of work, an author should do before asking us to read something?
The Scholar’s Telescope
In The Scholar’s Telescope, John McGinnis (Northwestern Law) argues in favor of AI-assisted legal scholarship. He argues, reasonably enough, that the governing purpose of legal scholarship is to “improve warranted public understanding of law,” and that this purpose, rather than the quantity of unaided labor an author invested, should determine how AI may be used. “Unaided labor is a means, not the end.”
He thinks it’s fine for legal scholars to use AI extensively, provided they verify their material claims, investigate and acknowledge their intellectual debts, and understand and stand behind the judgments they publish. He is careful to note that these are conditions of scholarly integrity. Whether a work deserves publication depends on the traditional criteria such as originality, importance, and illumination.
McGinnis’s argument sounds logical when he says “a mature scholar does not improve an article merely by making it harder to produce.” His telescope analogy makes the point: astronomy gains from instruments that extend what researchers can see. Scholarship can likewise gain from instruments that extend what researchers can investigate, provided they know how to supervise them.
AI Slop
In AI Slop, Jessica Silbey and Woodrow Hartzog (both at Boston University Law School) try to pin down what exactly constitutes AI slop before explaining the ways it is corrosive to things we value.
Their core argument straddles the line between definition and taxonomy. In their words, “AI slop is any output of a generative probabilistic automated system produced with little exertion that asymmetrically burdens recipients and tends to degrade cultural domains.” But although they provide a definition, they also describe a spectrum: outputs can be more or less “sloppy” according to the relative presence of three constituent elements: negligible exertion, asymmetrical imposition, and domain degradation.
Given their definition, it is hard to disagree with anything Silbey & Hartzog actually say about why slop is bad. If slop is that which “burdens recipients and tends to degrade cultural domains,” it’s not surprising it’s bad. They explicitly acknowledge that their definition is normative. The harder question is how much it helps us decide whether any given use of AI will produce those bad outcomes.
Consider their beneficial examples. AI-generated reports of earth tremors, they say, satisfy the first two elements of their definition: the output takes no human effort to generate, while its recipients expend time and cognitive resources evaluating it. But these reports can improve earth science rather than degrade it. Phew! Live translation at low-stakes gatherings gets a similar pass. Elsewhere, they recognize that artists and writers can use AI for brainstorming or assistance while still engaging in creative, communicative practices.
But Silbey & Hartzog are cagey about how much assistance they think is acceptable, and why? What distinguishes productive assistance from the offloading of effort they condemn? If “domain degradation” decides the question, how do we identify it before we have already decided whether the use is good or bad? Also, domain degradation is necessarily a problem of cumulative effects, which individual authors might not be in a good place to assess.
More useful to me than the definition is their taxonomy of the harms:
- Slop crowds out human expression, making desirable work harder to find and eroding presumptions of trust and sincerity.
- It cheapens human effort, including the care, learning, and relationships that effort can sustain.
- It encourages fraud, indifference, unearned confidence, and sloth, while introducing errors that can compound over time.
Crowding out
For a while, my spam filters seemed to have made my inbox manageable. Now my email is peppered with highly articulate, individualized requests for “just 20 minutes of your time.” I have started using AI to triage my personal email account, which means that bots are writing emails and reading them.
McGinnis proposes using AI to check sources, quotations, and other matters of integrity, while leaving judgments of scholarly merit with human editors. That’s all well and good, but it doesn’t respond to the problem of the scale of the demand being put on human readers.
Many of our institutions have systems and processes built around a historical experience of a certain volume of production. A sudden increase in that volume, may overwhelm them, even if some of it is genuinely good. But if the new abundance of scholarly production is predominantly low quality, then we have two problems rolled into one.
Cheapening
Silbey & Hartzog argue that a lack of effort cheapens author and reader alike, but I think the value of effort is highly context specific. Some writing matters partly because somebody bothered to write it. A personal letter can communicate care through the time and attention its author invested. Time and energy are scarce resources; so spending them is a way of making a statement credible.
Silbey & Hartzog don’t put it in these terms, but I see the beginnings of a Market for Lemons problem here. Recipients who can’t distinguish thoughtful communication from simulated thoughtfulness will be inclined to discount both. That makes genuine effort less rewarding and may eventually discourage it. If this really is a market for lemons, then everything unravels.
Beyond signals, Silbey & Hartzog also value what effort does for the person exerting it: developing skills and judgment, building interpersonal relationships, and more. This is undoubtedly true, but again, highly context specific. Mechanical intervention is a good idea if you are lifting heavy things at work, but self-defeating in the gym. Silbey & Hartzog don’t have much to say about domain variation and even less to say about how to balance the costs and benefits of reduced effort. McGinnis is on firmer, narrower ground here because his point is that, in scholarship, the public contribution should ordinarily take priority over a preference for unaided production.
Slop as a vector for vice and error
Silbey & Hartzog also see slop as a vehicle for moral and cognitive offloading. People can blame the machine for mistakes or take credit for work they didn’t do. Plausible errors can enter the record and then propagate into the next round of production.
Who pays for the reading?
The papers are not entirely at odds, but they disagree about how much the production process should count. Silbey & Hartzog commend policies protecting human authorship and requiring disclosure of substantive AI use. McGinnis generally resists such disclosure requirements and treats unaided production as valuable when it serves some further scholarly purpose.
I am closer to McGinnis in theory and to Silbey & Hartzog in practice. Silbey & Hartzog call negligible effort the “original sin” of AI slop. But other than being a seismologist, it’s not clear what exactly an author has to do to be regarded as investing more than negligible effort. If an author puts substantial thought into the research and then carefully edits an AI-assisted draft, is that enough? Or perhaps the better question is: when is that enough, and when isn’t it?
McGinnis, on the other hand, leaves me unconvinced about the cost of making his responsibility regime work. He acknowledges that abundance “can exhaust editorial attention and make genuine originality harder to find,” then proposes automated integrity screening and human assessment of merit. That is a reasonable division of labor, but it still leaves the expensive part with the humans.
His answer also depends on authors taking their responsibilities seriously. In theory, certifications and sanctions might help, but it would be nice to see a positive case study where they actually have.
Perhaps we are in a transitional moment, and AI-assisted scholarship is about to become reliably awesome. But it’s very far from it in domains like legal scholarship.
Reading AI writing
I research at the intersection of law and AI, and I use AI a lot. Claude Code and Codex have improved my life in ways large and small. I’m very pro-AI, up to a point.
And bad writing is that point.
I read Silbey & Hartzog’s paper from start to finish in one session because it is a well-written, well-argued piece on a topic that interests me. On my first pass through McGinnis’s paper, I stopped well short of the end. By halfway through, I had a strong impression that I was reading extensively AI-generated prose and a diminishing expectation that more time would repay the effort. To be fair, I have no firsthand knowledge of how this paper came to be, and I assume that McGinnis at the very least supplied the key insights and organizing ideas.
I am not opposed to AI writing. Some of this post is, or was, AI writing. I honestly couldn’t tell you which phrases began with notes I dictated (using yet more AI), which came from Claude or ChatGPT, or which were edited and restated until I am pretty sure anyone who is not a zealot would agree that they are, in every relevant sense, mine.
I would not say that that Claude and ChatGPT are bad writers, as such. But I read a lot of their output and their familiar habits become more grating with repetition. Here are seven examples from McGinnis, with page references:
- “The three words in the formulation supply the architecture.” (p. 3.) This is a grand announcement of an explanation the reader is about to receive anyway.
- “the best analogy for AI is not the ghostwriter or the cheat sheet. It is the telescope, the microscope, the database, the statistical package, the search engine…” (p. 4.) Not x but y, and also the telescope is a useful analogy. The expanding parade of instruments adds less with each arrival.
- “The relevant line, however, is not between a ‘tool’ and a ‘non-tool.’ It is between tools that require different kinds and degrees of supervision.” (p. 5.) Not x but y. The qualification is fair enough. The reset into another corrective contrast is already becoming familiar.
- “Truth-seeking is indispensable to that enterprise, but it does not exhaust it” (p. 3), followed by “Truth is indispensable, but it is not exhaustive.” (p. 8.) Not x but y repeated! The second formulation sounds like a fresh distinction even though the first has already made it.
- “An empirical article may do so by discovering a fact. A historical article may recover a doctrine’s genealogy. A doctrinal article may show that apparently unrelated cases rest on a common principle. A critical article may expose an assumption…” (p. 8.) Normative and conceptual articles follow. The examples differ, but the six-part enumeration painfully prolongs this point.
- “The lesson from this literature is not that scholarship has no purpose. It is that scholarship’s many immediate functions must be distinguished from its governing public justification.” (p. 8.) We return to the central distinction in another familiar arrangement.
- “The point is not ritual human involvement. It is a process reasonably capable of uncovering the errors the use is likely to create.” (p. 24.) Again, the same delivery.
Of course, contrasts and parallel structures can be useful. But the cumulative experience is dire. Too many paragraphs announce, enumerate, or reformulate a proposition before the argument moves on. The reader has to keep checking whether a new formulation contains a new thought. Far too often it doesn’t.
The clichés and stylistic tics are part of the problem, but the larger problem is informational. With this kind of AI writing, I often feel that the payoff declines as I read: early paragraphs establish a point, and later ones keep finding new ways to say it. In my experience, the decline in informational content happens across the document as a whole, but also within many individual paragraphs. Sometimes it’s helpful to elaborate on the point that you’ve just made, but over-elaboration is tedious.
Why is AI writing so boring? I suspect this has something to do the mechanics of next token prediction, but it’s hard to pin down the exact mechanism. I have been teaching at law schools for over 20 years, and I have graded my fair share of badly written papers. Bad AI writing makes me bored much faster than bad human writing. It signals to my brain, at least, “you’re not going to get much more out of this.”
It doesn’t have to be this way.
Better writing with AI
Much of what follows is consistent with McGinnis’s verification requirements and Silbey & Hartzog’s concerns about offloading work. What I want to add is a practical (and non-exhaustive) account of the editing that should happen before an AI-assisted document reaches its readers.
Understand your role
Exertion comes in different forms. Consider a doctor who explains her clinical reasoning aloud during a consultation, then reviews and edits an AI scribe’s draft. Has she avoided the cognitive work, or found another way to do it?
Silbey & Hartzog consider this problem directly. Drawing on Helen Ouyang’s account of using an AI scribe, they worry that reviewing an already composed note may not require the thinking involved in writing one. That is a serious but only partial objection. Patient records serve patients, other healthcare providers, and insurers, and clinicians usually have to write separately for each one. This sounds like exactly the kind of translation task that generative AI is really good at, most of the time. Obviously, there are risks that will be measured in human lives, but so too are the rewards. The key to mitigating those risks is not to tell doctors to do more work, but to help them understand how the nature of that work has changed. In an insightful account, Sari Altschuler and her coauthors develop the editorial role in Clinician as Editor: Notes in the Era of AI Scribes, 404 The Lancet 2154–2155 (2024). They emphasize narrative and editorial training, organizational support, and the cognitive work that still has to happen in some form. In short, doctors need training and structures to help them avoid sloppiness, and understanding the evils of AI slop is really only the starting point.
Use AI to reduce friction
For a recent article, I used Claude Code to create a database of appellate-level fair use decisions and query them with seven specific questions. The resulting Excel table helped me identify examples of pathologies I was planning to write about.
I was familiar with almost all of the cases and I still read the relevant extracts. But I would not have painstakingly scanned hundreds of decisions and made notes in a table for an analysis that was far from the article’s main point. Even if I once might have undertaken that heroic effort, now my time is better spent doing other things.
This is close to the use McGinnis describes: let the machine help locate and organize the material, then read and evaluate.
Use AI to add the right kind of friction
Silbey & Hartzog worry that slop imposes an asymmetric burden on recipients. One response is to take more of that burden back before pressing send. I would amend the familiar maxim to read: “If you can’t be bothered to edit it, why should I bother to read it?”
AI can help with editing if you ask questions designed to elicit criticism.
ChatGPT came up with:
- Which paragraphs add no new claim, evidence, or necessary qualification?
- What would the author I am criticizing say I have overlooked?
- What does the reader learn from this section that they did not know at its beginning?
I didn’t particularly like any of those, but I leave it for the reader to judge. Here are some of my favorites:
- Predict all of the outraged hot takes I am likely to see online once I publish this paper.
- Critique this paper as a [Economist, statistician, pretentious cultural critic, jaded internet law scholar, computational linguist, computer scientist, industry insider, …]
- You are a second-year law student screening articles for a top-tier journal …
- If this paper occupies one point along a spectrum of reasonable opinions on this topic, critique it from the opposite side of that continuum.
- Highlight for me seven or eight sentences and phrases that, on reflection, sound a little bit like an AI wrote them [optional: and suggest how they could be changed].
- Explain why the attached paper makes no real contribution to the literature.
Some of these work better in a fresh window with the memory off.
Look beyond the model for validity
Hundreds of legal decisions have addressed lawyers’ copy-pasting hallucinated citations into court documents. Damien Charlotin’s database catalogs decisions from the United States and overseas. You might think that after the first few well-publicized incidents, lawyers, at least would have got the memo. But perhaps it’s not that surprising because when a system’s answers are fluent and often correct, it is easy to misunderstand what is actually going on under the hood or to simply become complacent.
Training a base language model to predict the next token rewards it for learning patterns in its training data. Those patterns can encode a great deal of knowledge. But this process does not supply any guarantee that the model has learned accurately, or that the material it learned from was true in the first place. Post-training and other bells and whistles make the models more useful, but they don’t supply any additional guarantee of the truth of the outputs.
So, if you are using LLMs you need a source of validity beyond the model. Some, but not all, coding errors announce themselves when the program fails, so that’s helpful. But in every other context, just because something looks right doesn’t mean it’s right.
The world would be a very dull place if we didn’t make mistakes, but we should try to make them our mistakes.