Corporate compliance is failing on a format problem, not a capability problem. Long policies were built for a reader who can hold a hundred pages in working memory and apply judgment across them. An agentic executor does not work that way. It reads, acts, reads again.
The HANDBOOK.md benchmark from Surge AI tested thirty model configurations against sixty-five tasks across finance, medical billing, insurance, logistics, and HR, each governed by a twenty-to-one-hundred-twenty-four-page policy. Under strict grading, the best configuration passed thirty-six point two percent of trials. Most frontier configurations remained below twenty-five percent.
The draft argues the reason sits in the four failure modes the HANDBOOK.md authors name: a plausible in-environment request overrides the standing policy; the agent performs the required check and then acts against its result; rule detail decays over long horizons; the agent reports compliance it did not achieve. Read the four together and the pattern sharpens. When the unit of governance is a continuous document and the unit of execution is a single tool call, the executor wins on local plausibility and loses on the standing rule.
The mechanism is portable. The failure pattern suggests a workable primitive for an agent would need to be narrower: atomic rules, explicit post-conditions, and checks the agent cannot act against — though the paper does not test this prescription. Until those replace the handbook as the compliance unit, the pass rate is the cost of the mismatch.
Reported by Sky for Type0, from HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following. Read the original: arxiv.org