A memory backed search method called FLEET, posted as a single author preprint, lets a 3 billion parameter model keep notes on its prior attempts and steer away from dead ends — with reproducible code behind it.
A 3-billion-parameter language model just matched its own best answers using about half the guesses, by keeping notes on what it has already tried.
The method, called FLEET and posted as a preprint on 23 September 2026, replaces memoryless trial-and-error with a memory-backed search. Instead of treating each attempt as independent, the system attributes each attempt's reward to specific tokens and uses a modified tree search to steer the next try away from dead ends. The result, according to the GitHub repository: on the LiveCodeBench v6 easy split, Llama 3.2 3B Pass@32 climbs from 0.6 to 0.66 — a roughly 10% lift — while GSM8K matches the prior best at about half the iterations and gains 7 solved problems.
The catch is the size of the evidence. The numbers come from a single author, Oleksandra Vitko, via Reddit self-promotion of a 25-page preprint, not a peer-reviewed paper. The benchmark is narrow: a 3B model on the easy splits of two tests. The three sources — Reddit, the arXiv abstract, and the GitHub README — agree on the headline numbers but vary slightly on the exact figure (0.59 to 0.69, 59.9% to 66.2%, 0.6 to 0.66). Independent reproduction on larger models and harder splits is the obvious next test.
The code is available today — pip install fleet-search — so the falsifier is the next run, not the next press release.