BEAM, peer reviewed at ICLR 2026, scores whether an agent still trusts the right fact across very long conversations, tickets, and documents. past.dev's run is the first on the open harness.
A new memory layer for AI agents does three things at once: it tracks which fact is still true today, what it replaced, and who is allowed to see it. BEAM, the peer-reviewed long-term memory benchmark published at ICLR 2026, is the public scoreboard where that capability gets measured, and past.dev's first public run on the open harness leads the board at every history size the benchmark tests.
BEAM, described in its paper on arXiv, asks a sharper question than a similarity search. Instead of asking an agent to find text that looks similar, it asks whether the agent still trusts the right fact after the underlying reality has changed: a pricing tier renamed, a teammate who left, a feature flag that flipped. To stress that distinction, BEAM runs out to 10 million tokens, well beyond the native context window of any frontier model, where the questions are 200 spread across 10 long conversations, ticket threads, and document histories.
The leaderboard split, as of past.dev's public methodology page and its open evaluation harness on GitHub, looks like this. At 10 million tokens, past.dev's system scores 85.03% versus a prior best of 68.0% published on BEAM, a 17-point gap. At 1 million tokens, it scores 90.65% versus 79.1%. At 500,000 tokens, 89.63% versus 80.1%. At 100,000 tokens, 92.08% versus 86.2%. Past.dev published all four runs on September 29, 2026, using the public harness and ExaBase's published BEAM prompts as the judge, so the rows are meant to be re-run by anyone with the time.
The number that should sit alongside the leaderboard is the one in the BEAM paper itself. When models are simply handed the full history, the best score is 25.9% at 1 million tokens. That gap, from 25.9% raw context to 90.65% with structured memory, is the size of the problem the category is solving. Long context is not enough. An agent has to know which fact is current.
Per-asker access control sits on top of the timestamp store. The system retrieves the current fact and returns it, but only to users whose permissions match. In the press release, past.dev frames this as deciding what a given user can see "without leaking a private email" they should not have access to. That permissioning is what turns a memory layer into something a multi-tenant enterprise can deploy.
Past.dev has open-sourced the BEAM evaluation harness, kept the judge prompts unchanged, and published the methodology page that acknowledges the apples-to-oranges problem: "models, judges and configurations differ between rows." Independent reproduction on a non-vendor harness would strengthen the leaderboard claim, and the harness is there to be run. Today's public evidence is past.dev's own published run plus the open harness code, not a third-party leaderboard.
Two things to keep separate. The "previous best" rows above are the prior best published on BEAM, not an industry-wide state of the art across every long-context memory system. And past.dev's claim that its system runs in production for large teams across very large volumes of data is unnamed and unverified by this release; the leaderboard is about the benchmark, not the customer count.
Past.dev's Frontier Memory ships via API, supports the Model Context Protocol for tool-using agents, can be self-hosted, and is offered as a dedicated region on Enterprise. The company lists SOC 2 Type II, ISO 27001, and ISO 27701 certifications, and says no customer data is used to train models. None of that is the scoreboard; it is the deployment surface a reader has to check next.
The next BEAM numbers will land on the open harness past.dev has already released on GitHub. Readers will have a cleaner question to ask by then: not whether the agent remembers, but whether it still trusts the right fact, and whether a third party can re-run the number to confirm. Past.dev has put the harness on the table. The next run will show whether 85.03% is the new floor for the leaderboard or just the first data point published.