ICML 2026oral0 citations

Position: Stop Automating Peer Review Without Rigorous Evaluation

Joachim Baumann, Jiaxin Pei, Sanmi Koyejo, Dirk Hovy

Abstract

Large language models offer a tempting solution to address the peer review crisis. This position paper argues that **today's AI systems should not be used to produce paper reviews**. We ground this positing in an empirical comparison of human- versus AI-generated ICLR 2026 reviews and an evaluation of the effect of automated paper rewriting on different AI reviewers. We identify two critical issues: 1) AI reviewers exhibit a *hivemind effect* of excessive agreement within and across papers that reduces perspective diversity. 2) AI review scores are trivially gameable through *paper laundering*: prompting an LLM to rewrite a paper could significantly increase the scores from AI reviewers, demonstrating that LLM reviewers are easy to game through stylistic changes rather than scientific results. However, non-gameability and review diversity are *necessary but not sufficient* conditions for automation. We argue that **addressing the peer review crisis requires a science of peer review automation**---not general-purpose LLMs deployed without rigorous evaluation.

LLMBenchmark
BibTeX
@inproceedings{icml2026_positionstopauto,
  title = {Position: Stop Automating Peer Review Without Rigorous Evaluation},
  author = {Joachim Baumann and Jiaxin Pei and Sanmi Koyejo and Dirk Hovy},
  booktitle = {ICML 2026},
  year = {2026}
}