For two years, the long-context research community has treated post-trained linear attention as the most promising path off quadratic cost. A new arXiv preprint from Alexia Jolicoeur-Martineau, Rhea Sanjay Sukthanker, Pashmina Cameron, and Emy Gervais suggests the comparison was rigged from the start.
On Needle-in-a-Haystack and BABILong, the paper reports sliding-window attention landing 2 to 10 times higher than linear attention — not a marginal win. The configuration is mundane: a 64-token window plus four attention sinks, no post-training, low memory at inference. What changes is the baseline. Linear-attention variants from SUPRA, Hedgehog, LoLCATs, Liger-GLA, MOHAWK, Mamba-in-the-Llama, DiJiang, ARWKV, Llamba, QLinAtt, QRWKV6, and QRWKV7 had been measured against standard attention, not against the simpler, deployable fix. The authors' MMLU-5shot recovery table tells the same story: SWA at 93.2, Llamba at 91.5, QRWKV6 at 92.4, QRWKV7 at 86.4, DiJiang at 88.7, ARWKV at 84.1, Mamba-in-the-Llama at 67.7.
Two of the four authors work in Microsoft's Applied Sciences Group, and the paper has not been peer-reviewed. The 2–10× claim is scoped to those two benchmarks, not universal across long-context evaluations. Independent replication has not yet arrived. Still, the methodological point is hard to dismiss: a simple pattern hiding in plain sight outperformed a research line that has consumed real post-training compute.
The paper's recommendation is blunt. Practitioners running long-context inference should switch to sliding-window attention rather than continue investing in post-training pipelines to linear attention — or accept that linear attention, to compete, must be trained from scratch with extensive post-training. The field's compute bill may have been chasing a phantom.