By keeping object identity and permission attached to each executed step, the FlowReview framework cut over blocking from 86% to 0% in controlled multi agent experiments with no authorized work loss.
When multiple AI helpers work the same task and hand evidence to one another, a request that looks fine in isolation can become dangerous once two innocuous pieces of information from two different agents get combined. The merged action crosses a line neither agent would have crossed alone. That is the hard version of multi-agent safety, and it is the part current guardrails handle worst.
Today's safety systems read each step on its own. The result is a tradeoff without a comfortable middle: either let too many dangerous combinations through, or refuse so many legitimate actions that the collaboration stops being useful. A new arXiv preprint, "Deny Without Disabling: Authorization-Paired Evaluation and Control for Multi-Agent Systems", argues there is a third option. The fix is not a smarter filter. It is a different design for what the system keeps track of while the agents work.
The paper calls its design "authorization-paired evaluation." The mechanism is plain. A safety setup should not just record that a tool was called or a file was opened. It should keep the identity of the object and the permission attached to the execution step that produced it, and verify that the link is still intact when a later agent reads or acts on that information. In other words, who this piece of data is and what the original agent was allowed to do with it has to travel with the data through the whole composition.
The authors turn this into a framework called FlowReview, released alongside the paper. FlowReview stitches three pieces into one path: object resolution, which follows the identity of a piece of information across handoffs; permission ranking, which keeps track of which agents are allowed to do what with it; and deterministic enforcement, which checks that the two still match before an action is allowed to commit. The project page walks through how the same identity and permission handle follows a piece of data as multiple agents contribute to a task.
The headline number comes from the paper's own controlled composition experiments. When the system had to read the combined information flow before allowing a commit, the share of legitimate actions that got wrongly refused, the "denied-commit rate," fell from 86.0% to 0%. Authorized supply, the work that was supposed to get through, did not drop. That is a striking result on its own terms, and it is the part the field has been waiting for: a way to cut over-blocking without paying for it in missed work.
Preserving the information and preserving the lineage of where it came from are not enough on their own. If the system cannot keep identity and permission linked to execution, composed agents end up acting on data whose original permission was lost in the handoff. The negative finding is the mechanism's strongest evidence: the missing piece really is the link between object identity and the executed step, not better record-keeping in general.
The result is from controlled composition experiments on the authors' own setup, not from a deployed multi-agent system in the wild. The work is a single arXiv preprint, not peer-reviewed, and no independent replication has been published. The framework asks for components whose outputs can be verified, which is an architectural commitment a deployed system has to actually make, and most current agent stacks are not designed that way yet. Anyone reading the 86.0%-to-0% number as "multi-agent AI is now safe" is reading past the caveat.
The right safety design for collaborative AI is not to forbid collaboration. It is to keep object identity and permission attached to execution, through components whose outputs can be verified, so composed flows get reviewed before they commit. That is what makes more capable multi-agent systems possible to build without giving up the line.
The next question is whether the same mechanism holds up outside the authors' setup. Code is public at the FlowReview repository; whether other groups can reproduce the 0% denied-commit rate on independent compositions is the watch item for the next paper.