Five leading AI model families, ranked together, closed 60% of the gap to where real unpaid web readers actually highlight and beat the best single model in an independent replication.
The leading open-source prompt compressor, LLMLingua-2, landed below naive sentence truncation in a 120-document test against real unpaid web readers. It was statistically indistinguishable from random selection.
Fusing rankings from five different frontier model families, plus a simple position prior, did the opposite. The unweighted fusion closed 60% of the gap to a human-reader ceiling and beat the best single model by +0.0159 average-precision points (Holm-adjusted p=0.019). That edge survived a pre-registered replication on 217 independent documents at +0.0179 (p=0.042).
Lead-bias and length features recover only about 5% of the gap, so the missing signal is semantic, not positional. Kazuki Nakayashiki and Keisuke Watanabe posted the preprint to arXiv on 3 August 2026 with the replication already attached. The work is single-source, not peer-reviewed, and the model identifiers (GPT-5.4, GPT-5.5, Claude Sonnet 4.5, Gemini 3.1 Pro, Gemini 3.6 Flash) come from the authors' own test harness.
A distilled 8-billion-parameter open-weight student kept 90% of the fusion's edge, reaching statistical parity with the strongest single frontier model. A local-context student kept only 63%. Model diversity, not model size, is the cheapest known upgrade for attention prediction.