Anthropic's new top tier assistant nearly matches the previous leader at about half the price, but Zvi Mowshowitz's long test suggests it works best as part of a larger system, not a daily driver.
Anthropic shipped Claude Opus 5 this week with a price-and-performance pitch: a new top-tier assistant that nearly matches the previous leader at about half the per-token API price. The trade is what the model is like to spend a day with.
Independent reviewer Zvi Mowshowitz spent a long week with Opus 5 and concluded the model is built to run as a piece of a larger system, not the one you talk to all day. The review cites repeated verbal tics, overlong sentences, conversational loops that turn confrontational, and tangents into paranoia. Mowshowitz's verdict: a strong subagent and a poor daily driver, with the same complaints showing up in Opus 4.7 and 4.8.
On the numbers, Opus 5 is still near the top of the field. On Vals's independent index, it ranks #2 at 74.82%. The Artificial Analysis Intelligence Index puts Opus 5 in the lead at 61 points. Anthropic's own launch post positions the model as suited to local thinking, less so to long-horizon planning, and warns that higher effort settings can send it spinning. Medium effort is usually enough.
What remains unknown is how much the user-vibe complaints generalize. Mowshowitz's read is one informed reviewer's long-week test, not a population survey. The next signal: whether the argumentative loop shows up when Opus 5 is used as a worker in a multi-agent harness, the way Anthropic's pitch suggests, or only when it runs the show.