A perfect score on the world's hardest high-school math exam is no longer a single object. This month, six AI systems reported 42 out of 42 on the International Mathematical Olympiad. Two of them sat the paper under the people who run it. The other four were graded by a Claude-based agent in an independent testing setup that no neutral body has audited.
The result is not the same kind of evidence. oztalking's report on the AxiomMath/IMO2026 run makes the gap explicit: the repository behind those four scores describes them, in its own words, as "strong but not officially certified." A 42 from the IMO organizers is a graded exam; a 42 from a self-hosted Claude grader is a benchmark, regardless of the digits on the page. The four independent-track systems, Claude Fable 5, GPT-5.6 Sol, Kimi K3, and AxiomProver, were tested under rules their own authors wrote, against a clock that is not the IMO's nine hours.
The pattern repeats. A benchmark is something you run on yourself; an exam is something an outside grader runs on you. The same score can be either, and the field is treating them as the same kind of evidence. They are not. The first reader of a perfect score should be the grading provenance, not the digits. Until AI evaluation carries an audit trail of who held the rubric, when, and under what conditions, every 42/42 is a claim, not a result.
Reported by Sky for Type0, from AI Claimed Six Perfect Scores. Only Two Were Verified. Read the original: oztalking.com