The first time a benchmark number crosses into general press, it usually marks a shift in what the field measures. Model IQ has given way to agent follow-through: who can take a real-world job, run it across tools and files, and hand back a verified deliverable. That move is correct. The new problem is who certifies the score.
When the company whose agent tops a leaderboard also publishes the full evaluation pipeline on its own GitHub, the number stops being a cross-vendor ranking and starts being the vendor's process confidence, made public. Baidu's claimed 94.6% peak on PinchBench v2, dated July 17, is the first clean instance. The pipeline at Baidu-AI-Search/PinchBench-Evaluation ships with 147 task snapshots, deliverables, and LLM judge rationales. Auditable, yes; auditable against a definition Baidu itself defined.
The structural issue runs deeper. PinchBench scores a model together with its agent framework, so any vendor that tunes its framework against the test quietly tunes the score. Wire copy will treat 94.6% as a victory. The portable reading is that Baidu is the most confident vendor in its own evaluation loop, with Claude, Qwen, and GPT entries measured on the same surface. The next agent-benchmark headline, from any vendor, deserves the same three-question check: who publishes the test, who grades the runs, and who maintains the pipeline.
Reported by Sky for Type0, from 百度文心助手任务Agent登顶国际权威榜单,超越Claude、GPT拿下全球智能体冠军. Read the original: qbitai.com