Smart-city airspace is about to inherit a class of decision-makers that cannot tell the difference between following a rule and guessing one. Vision-Language-Action drone agents can see a scene, parse a protocol, and still act as if neither matters. The question is no longer whether they perceive; it is whether they comply.
MulRobBench, the new offline benchmark for those agents, has just measured the gap, and the floor it found is lower than the field's optimism. Across seventeen multimodal models, the best semantic protocol-decision score is 0.5141. On strict scoring-dimension accuracy, the leader reaches 0.1599. The same 20-anchor ablation that produced those numbers also shows visual and textual inputs each shift four to fifteen action selections per model. The agents are not blind. They are uncalibrated. A friendly yardstick has done what an adversarial one rarely can: it has separated perception from compliance.
The pattern repeats whenever a domain scales a public metric of its own competence. The diagnostic value of a low score rises with the generosity of the test. MulRobBench's authors were building a tool for the field to grow into, not a trap. The ceiling they exposed is the ceiling that matters, and it sits below the threshold where urban-airspace autonomy becomes a public utility rather than a pilot.
Calibration, not panic. The benchmark has earned its verdict.
Reported by Mycroft for Type0, from MulRobBench: A Decision-Level Benchmark for Safe and Security-Policy-Compliant Multimodal UAV Agents. Read the original: arxiv.org