A new agent harness splits "the model said so" into five testable steps. Its own pilot didn't show it was better.
A model can propose an action without being authorized to run it. Running it is a different claim from it having the intended effect in the world. A new arXiv preprint, From Proposal to Verified Effect: Praxa, an Evidence-Bound Harness for Governed AI Agent Execution, turns that gap into the central object of design. The paper is the first to write the proposal-to-effect pipeline as architecture rather than as rhetoric, and it is unusual in the agent literature for stating, in the same preprint, what its own pilot did not show.
Praxa is an "agent harness" — the system around a model, not the model itself. The authors, who identify as praxa-labs, name five states an agent action can occupy: proposal, authority, dispatch, verified external effect, and serving promotion. Each transition between those states is meant to be testable on its own, and the language is meant to be borrowed. The point is not that Praxa alone closes the gap, but that the field can now name the gap in five pieces and argue about each one.
Four components handle the transitions, as described in the full paper. Deterministic admission uses rules to decide whether a proposal is even considered. Brokered execution routes the proposed action through an intermediary that decides whether to run it. External read-back checks, with something other than the model, what actually happened in the world. Reconciliation compares that read-back to the original intent, and reviewed promotion gates any answer before it reaches a user or downstream system. The four names are the contribution: they map the path from "the model said so" to "this is what happened" onto code that can be inspected, instrumented, and replayed. The accompanying GitHub release, preprint-v1.4.0, puts the harness and the four evidence lanes under a public repository for inspection.
The results are the more honest half. The authors run four "evidence lanes" themselves. There is no independent reproduction, no third-party benchmark, and no production deployment, and they say so.
Lane 1 is a repo-local audit: the harness passed 1027 of 1027 unit tests and 89 of 89 Workerd tests (Cloudflare's JavaScript runtime, used for parts of the dispatch layer), and met four coverage floors across 363 source files. Raw per-test transcripts are not included; readers get pass counts and floors, not diagnostic logs.
Lane 2 is a provider-backed pilot on Terminal-Bench Core 0.1.1, a 12-task agent benchmark, run on curated tasks. Baseline and the reliability layer each passed 17 of 36 strict trials — a tie. The reliability layer used 37.49% more input tokens and 50.73% more output tokens. The authors state plainly that the pilot does not support superiority.
Lane 3 is a coordination-proxy comparison run after a debug pass. Baseline and source-authored candidate each completed 180 of 180 trials at equal accuracy, with full hermetic (fully isolated, no network or external state) crash recovery and zero protected-attribute violations. The candidate used 37.11% fewer tokens, an estimated 33.84% lower endpoint cost, and 11.63% fewer steps. The authors flag that this is a post-debug, source-authored, two-order setup, and that it does not establish improved quality, latency, or production behavior.
Lane 4 is a deployed source and config review: bounded reflection, recall accounting, memory compilation, and tool-health paths are present in the codebase. There is no production outcome lift.
What the four components enable is a working vocabulary. The field can start to ask, of any agent claim, which of the five states is being asserted and what evidence backs that specific state rather than a downstream one. A team that learns to publish its null results next to its architecture is practicing the kind of evidence discipline real capability and real deployment will eventually need. Praxa is one team's attempt; whether the wider agent community borrows the vocabulary is the next test.
The authors also list what they are not claiming. The paper does not assert adversarial security, production safety, general specialist superiority, autonomous recursive optimization, or user benefit. That list, published in the same preprint, is the part most worth carrying into the next agent story a reader opens.