On a new benchmark across three AI research systems, runtime selection lifted scores while added contracts and verification dragged them back down.
A new benchmark asked whether AI research agents, the software systems that string together large language model calls to do pieces of the scientific process like searching literature, running experiments, and drafting papers, work better when they pick collaborators on the fly. The answer is yes to runtime selection, and a clear no to piling contracts and verification on top of it, at least on the systems tested.
The paper, "Can AI Scientists Coordinate at Runtime?" by Xisen Wang and posted to arXiv on October 1, 2026, frames an architectural question that has split the multi-agent AI-scientist community for the past two years: should these systems follow a fixed playbook, or should they coordinate and divide labor while running, the way human researchers do?
Most multi-agent research systems today use design-time orchestration. The agents and their hand-offs are decided before a run starts. The Wang paper proposes what it calls Runtime Agent Coordination, or RAC, in which agents are selected, routed, and revised during the run itself, against scoped work contracts and artifact-grounded verification.
To test whether that overhead pays off, the authors ran the same benchmark, ResearchClawBench, across three existing hosts: Agent Laboratory, EvoScientist, and ARK. Each host keeps its own models, tools, and permissions. Four cumulative conditions isolate the contribution of each layer. The first, N0, is the host as written, with its native fixed workflow. The second, R1, adds SharedNet, a layer for runtime communication between agents. The third, R2, adds runtime routing so that agents are chosen while the run is in progress. The fourth, R3, layers on scoped work contracts plus artifact-grounded verification.
The headline result, the authors write, is that R2 yielded the highest observed mean score for each host. Adding contracts and verification in R3 reduced those means relative to R2, with the size and direction of the change varying by host.
The paper implements contracts and verification in a specific way: the verifier records its findings but does not block, retry, or roll back work. So when R3 dragged scores down, the mechanism is not that verification caught errors and made the system more rigorous. It is that an advisory layer spent part of the same fixed lifecycle budget without any way to stop bad work from continuing. The cost was real. The safeguard was not.
The code is public at systemind-team/Runtime-AI-Scientist, and the repository documents the four conditions, the host bridges, and a reproducible single-episode workflow.
The evaluation is single-seed and described by the authors as exploratory, not a head-to-head with statistical claims, so direction is what can be supported here, not effect size. The finding is also host-dependent: R2 helped across all three hosts, but R3's effect varied, which means the pattern may not transfer to other agent topologies such as coding agents or tool-use agents without separate testing.
Runtime routing, where the system picks agents mid-run, currently looks like the part of coordination that pays for itself in these hosts. Layered contracts and advisory verification, as defined in the paper, look like overhead in the same lifecycle budget. If a future version of verification can block or retry on its findings, that calculus might change. Until then, the cleanest design rule is to spend coordination budget on choosing who does the work, and to keep verification on a separate, actionable track.
The team says a follow-up with seeded evaluation runs and a blocking verifier is in progress.