When AI reviews AI, the verdict runs through judges that share training data, so agreement is correlation, not truth. That is the pattern underneath an arXiv benchmark of automated science this week, where FARS won by roughly 2x and the more interesting result was the dissent.
The benchmark, arXiv preprint 2607.28631, ran three frontier LLMs, Gemini, Claude, and GPT-5.4, over 75 papers scored on originality, rigor, clarity, and significance. FARS's papers landed at 2.14 to 2.47 on the 1 to 5 scale; the other three systems sat at 1.00 to 1.87. Gemini and Claude agreed with a Spearman rho of 0.907. GPT-5.4 correlated at roughly 0.32, well below the other two.
FARS also published the 15 research proposals that were used as the shared topics across all frameworks. The two highly aligned judges most likely share training priors, making their high agreement a correlation signal, not truth — GPT-5.4's dissent is the natural experiment that exposes this. The inference is portable: any time you stack sibling models as a panel, high agreement tells you the panel is correlated, not that the verdict is true.
The human editor's job moves up the stack: from grading each paper to weighting which AI judge to trust, and how to publish when the panel splits.
Reported by Sky for Type0, from Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review. Read the original: arxiv.org