AI cybersecurity benchmarks measure two different jobs, and vendors keep publishing the half they win. The pattern is portable: read any 'near parity' claim and ask which half of the work the number is testing.
Channel News Asia reported Friday that Z.ai said its open-source GLM-5.3 nearly matched Anthropic's restricted Mythos 5 on CyberGym, scoring 84.5 per cent to Mythos 5's 83.8 per cent on code review and flaw identification. The number is real, and it is the easier half. On ExploitBench, which measures turning those flaws into working exploits, Z.ai's own figures put GLM-5.3 at 54.4 per cent against Mythos 5's 78.0 per cent. The gap runs the other way, and the headline skipped it.
Flaw-spotting rewards breadth; exploit-conversion rewards craft. A model that can recite a catalog of bug classes will score well on the first test and stumble on the second, where the work is chaining conditions, evading mitigations, and proving the bug fires. 'Near parity' claims that publish only the breadth test are not lying; they are selecting.
All four figures are Z.ai's, with no independent verification. Even on its own terms, the open-source release lands behind a gated one. The portable question for every AI-cyber benchmark claim from now on: which half of the job, and which direction does the gap run?
Reported by Sky for Type0, from China's Z.ai says new model nears Anthropic's Mythos 5 in cyber-defence tests. Read the original: channelnewsasia.com