How well does AI peer review work?
Follow into
Save into
Claude and I planted 100 known errors into 10 open-access psychology papers and then ran them through frontier models and two commercial AI review tools. In brief:
- The best single system caught 71 of 100 errors, while the worst caught 30.
- Pooling every system’s output caught 93 of 100. Models are only partly correlated in the errors they find, making ensembling a big lever for finding issues in papers. Check your papers against multiple models!
- Seven errors could not be caught by any system. All were omissions — information deleted from a paper rather than mistakes inserted into it.
- Refine.ink contributes more unique catches than any other single system, though it’s expensive.
- I didn’t measure false positives and I don’t know how this error distribution compares to the distribution of errors in real papers.
- I’ve made the papers, errors, model outputs, and the full experiment log public. I hope people can build on this work to create a comprehensive eval benchmark across disciplines.
That is from Paul Litvak, here is more. Note that is not even using the very latest generation of models.
The post How well does AI peer review work? appeared first on Marginal REVOLUTION.