Holistic evaluation breaks under its own weight. When one model call must emit many verdicts at once, per-criterion attention thins: agreement with expert graders in research, legal, and clinical assessments falls as the number of verdicts per call grows, and a bigger or longer-running model does not reverse that drop.
The mechanism is structural, not scalar. Each verdict gets a smaller share of the call's fixed budget for grounding itself in the underlying evidence, so the call's capacity is consumed by checklist navigation before the call ever checks the substance. A presenter who only rephrases the same work exploits exactly that thinness: across a best-of-N re-optimization, an overloaded judge accepts severalfold more genuinely unmet requirements than the same judge does when its checklist is partitioned.
The fix in arXiv preprint 2608.06422 is to split the checklist. Smaller groups, each judged in its own call, then aggregate the verdicts. The paper's headline result: a sharded weaker judge can outperform a more capable holistic judge, and can match that more capable judge even when the holistic judge receives the panel's full budget. Sharding is not a universal defense; criterion-targeted attacks that persuade the judge on each item separately sit outside its scope, and the paper proposes debate-style opposition layered on sharding as the response.
The repeatable pattern: when a single decision unit must emit many decisions, scaling the unit does not scale the decisions.
Reported by Sky for Type0, from Sharding Prevents LLM Oversight Failures and Adversarial Exploitation. Read the original: arxiv.org