An arXiv preprint reports a 24.48% token cut on LongMemEval, a long horizon AI agent memory benchmark, at the same accuracy level, by dialing a single budget knob that ranks how memory slots are filled.
A new arXiv preprint reports a 24.48% token cut for long-horizon AI agents on the LongMemEval benchmark at parity accuracy, by dialing a single rank exponent that decides how much of the prompt budget goes to routine recall versus the rare states the agent is most likely to forget. The proposed controller, called Core-Tail World Model (CTWM), is the practical half of the paper's claim; the other half is a diagnosis of why current memory systems fail.
The paper, "Heavy-Tailed Memory Traces in Long-Horizon Language Agents", names the underlying pattern "heavy-tailed memory traces": under finite context and repeated retrieval, an agent's memory concentrates on a small core while rare events drift into a long tail where prediction errors pile up. The fix is rank-based. τ allocates budget by rank, not raw frequency, and keeps a summarized tail so rare states do not disappear.
The numbers come with caveats. CTWM is tested on Synthetic Graph World (5.9% token reduction, 13.6% lower bottom-half tail error) and ALFWorld, with LongMemEval as the closest thing to a real long-horizon test. The source is an arXiv preprint, not a peer-reviewed venue, and the lever is a single hyperparameter, the kind of knob that can reward benchmark-shaped priors and decay when agents meet real traffic.
The next question is whether τ survives on production traces, not just the synthetic ones the authors used to tune it.