Procurement teams are buying incident-response agents on the strength of vendor demos, and the field had no shared ruler to grade those demos against. ORCA-bench just published one. The 25.3% accuracy that the best frontier agent posts on realistic oncall root-cause tasks (and 10% on hard ones) is not a verdict on whether AI is ready for the pager; it is a floor. The ORCA-bench authors shipped a 50-gigabyte, six-day, expert-verified testbed and explicitly say real production is larger, more dynamic, and more idiosyncratic than the curated instrument they built, so the number is a lower bound, not a ceiling. The repeatable mechanism is short. Run your agent against the public test set. Strip source-code access and watch every metric collapse. Add it back. The gap that survives is the engineering work you actually have to fund, and the gap that disappears is the work the testbed was load-bearing for. Frontier model releases will not close the gap. Claude Fable 5, the strongest agent the ORCA-bench authors tested, did not. Whoever ships the next reliability agent now ships against a yardstick the field shares, and the procurement conversation moves from "can the model reason" to "what does it cost to close the measured gap on the public instrument." That is a better question, and it is the one the public test set makes possible.
Reported by Sky for Type0, from ORCA-bench: How Ready Are Language Model Agents for Oncall?. Read the original: arxiv.org