Any learned system whose features carry words inherits the same failure shape: the controller gets good at the language, the language gets changed, the score collapses. The pattern is small and it travels. A benchmark that holds wording steady is testing whether a learned controller has memorized lexical cues, not whether it can plan. A benchmark that rephrases is testing the plan. Same controller, two scoreboards, and the only thing that changed was surface phrasing.
The paper that makes this measurable is a benchmark from researchers at UT Austin and the University of Maryland. In a controlled test, a learned controller for agentic workflows clears 100% of held-out tasks, then drops to 75.9% once those same tasks are rephrased — the authors describe this as a bounded proof of concept, not live-LLM performance. The static playbook it was supposed to beat holds at 93.5% on the rephrased split. The controller costs 43% less on familiar wording and 49% less on the rephrased one. Both numbers describe the same economy. The gap identifies lexical generalization, rather than route execution, as the principal limitation — the paper's own framing, not an independent verdict.
The reader's filter is one question. The next time a headline says 'agentic AI beats static,' ask whether the held-out split changed the tasks or only the phrasing. If only the phrasing, the score is not a robustness claim. It is a memorization claim with a controller attached.
Reported by Sky for Type0, from Learning Compositional Meta-Routing for Agentic Workflows: An Executable Benchmark. Read the original: arxiv.org