ICML Study: AI Bans Don't Change Peer Review Outcomes — SkimNews

Get the Health newsletter
Daily health & science — research, biotech, public health, the studies worth knowing. Free.
- Microsoft Research scientists ran a randomized experiment at ICML 2026 in Seoul, with 17,886 reviewers assessing 24,661 papers under either an LLM-ban policy or a permissive policy allowing AI for understanding and polishing but not judging merits.
- Acceptance rates were nearly identical across policies—27% under the conservative ban versus 26.5% under the permissive policy—with average review scores of 3.31 and 3.32 out of 6.
- A post-conference survey of 1,486 reviewers found 22.5% of those told not to use AI admitted to using an LLM anyway—to brainstorm feedback, read submissions, and draft review text.
- AI text detector Pangram classified only 52.2% of reviews under the ban as fully human-written, compared with 37.0% under the permissive policy, suggesting self-reported non-compliance may understate the actual rate.
- Miro Dudík at Microsoft Research said the non-compliance figure exceeded his expectations, identifying heavy reviewing workloads and insufficiently clear rules as structural drivers of the violations.
- Kayvan Kousha at the University of Wolverhampton called the findings evidence that banning AI in peer review is "very difficult to enforce" because of the workload burden on academics.
Why it matters: With 17,886 reviewers handling 24,661 papers at one of the world's largest AI conferences, the experiment exposes a structural enforcement gap: workload pressures and unclear rules make AI prohibitions effectively symbolic. Conferences drafting guardrails now have evidence that ban policies neither altered reviewer decisions nor deterred roughly one in five reviewers—and Pangram's detection data hints the true non-compliance rate may be even higher.
Ask SkimNews




