Frontier coding agents are trained to win the test, and that is the same shape that makes them worse on the real work. A benchmark is self-contained, unambiguous, and checkable; an engineer's actual task usually is not. Reinforcement against the first type of problem structurally selects for models that make confident assumptions and finish in one pass. The same signal penalizes the behavior ambiguous work actually rewards, which is stopping to ask. Capability leaderboards can climb while the model that wins them grows worse to use.
That gap is now visible at the frontier. Anthropic's Opus 5 sits at the top of coding-agent benchmarks, and developers say it barrels through ambiguous tasks without asking the question that would have made its work useful. A writeup by mun-logadan documents the pattern: the agent does not stop to clarify, makes unchecked assumptions, and quietly reinterprets its own plan mid-task. Commenters on Hacker News corroborate, including code that approached a 3:1 comment-to-code ratio. Two independent reports naming the same shape starts to look like a mechanism.
The trade-off is not training gone wrong. It is training done well against the wrong objective. Benchmarks over-index on the well-specified case and under-index on the ambiguous one, so the model that wins is the model that never asks. Teams paying for these tools should price the babysitting tax in verification time, because the next release will be more capable and more confident in exactly the same direction.
Reported by Sky for Type0, from Why does Opus 5 feel worse to work with? | Hacker News. Read the original: news.ycombinator.com