Should a language model review your paper?

August 25, 2026 - 4 minute read -
research science research-ethics.md blog

Should a language model review your paper?

You submit a paper and then you wait. Weeks pass, then months. When the reviews finally land, one of them is three sentences long and the other clearly read a different paper than the one you wrote. Now turn the desk around. Your own inbox is holding four review requests, three from editors you’ve never met, and you’re the same tired person who owes everyone else a review by Friday. If either of those scenes made you wince, you already understand the problem this paper is chewing on.

The paper is a 2025 preprint, “AI and the Future of Academic Peer Review” by Sebastian Porsdam Mann and colleagues (arXiv:2509.14189). Its argument, said plainly and up front: a language model can responsibly take over narrow, supervised parts of peer review, and it must not take over the human judgment that makes a review worth reading.

The system is not lightly strained

The authors are blunt about the state of things. They describe “long publication delays, escalating reviewer burden concentrated on a small minority of scholars, inconsistent quality and low inter-reviewer agreement,” on top of the usual systematic biases. And the discouraging part: “decades of human-centered reforms have yielded only marginal improvements.” So the load really is piled on a few people, which is the second scene above, and the fixes we already tried barely moved it.

What the model is allowed to do

Here is the useful half. The paper argues that “targeted, supervised LLM assistance can plausibly improve error detection, timeliness, and reviewer workload without displacing human judgment.” Think of the mechanical jobs: checking whether a proof holds, flagging a citation that is missing, catching a statistic that cannot be right, pointing a reviewer at the three claims out of fifty that actually need a human to think hard. The authors go further and sketch fine-tuned, retrieval-augmented, and multi-agent systems that could make review “more reliable, auditable, and interdisciplinary.” Keep the word supervised in view. It is carrying most of the weight in that sentence.

What the model is not allowed to do

Now the half that should worry you. The authors list the failure modes as “hallucination, confidentiality, gaming, novelty recognition, and loss of trust,” and they refuse to file these under details to patch later. In their word, these problems are “constitutive” of what makes review legitimate, which makes them “governance choices as much as technical capacity.”

Two of them are worth stopping on. Confidentiality: your unpublished manuscript is a secret you handed to one editor, and pasting it into a commercial chatbot quietly breaks that promise for everyone who trusted you with it. And novelty recognition, which is my shorthand for the model faking competence: a language model is very good at producing a confident, fluent review of a paper it doesn’t actually understand, and a genuinely new idea is close to the one thing it never saw in training. A tired human at least half-knows when they are bluffing. The model does not know at all.

My take, with the caveats a preprint earns

Where do I land? The paper’s own recommendation is the sober one, and I agree with it: reject both “uncritical adoption [and] reflexive rejection,” and run instead “carefully scoped pilots with explicit evaluation metrics, transparency, and accountability.”

Two caveats before you quote it at your next lab meeting. First, this is a preprint. It has not itself been through peer review, which is a small irony worth saying out loud, so treat its framing as a well-argued position and not a settled result. Second, “supervised” is cheap to write and expensive to enforce. The thing that goes wrong is not a rogue AI seizing the journal. It is a swamped reviewer who lets the model draft the whole review and then skims the output, and no editor can see that from the outside.

Why you should care, on both sides of the desk

If you submit papers, the confidentiality problem is already yours, today, policy or no policy: someone in your review pile may be feeding your manuscript to a model right now. If you review papers, the honest question the paper hands you is which parts of your own reviewing are the mechanical parts a machine could take, and which parts are the judgment you’re willing to sign your name under. That line is the entire debate, compressed.

The paper doesn’t draw the line for you, and I’m not going to pretend I can either. But once you’ve waited most of a year for two careless reviews, “let the machine help, carefully” stops sounding reckless and starts sounding like the least bad option on the table. I might be wrong about how careful anyone will actually bother to be.