AI bug-finding is cleaving into two markets, and the tier where detection is actually won is not the one generating the headlines.
Semgrep's June benchmark put Zhipu's GLM 5.2 ahead of Claude Code on the same IDOR set, 39% to 32% at roughly $0.17 per vulnerability found. The same Semgrep test put their purpose-built multimodal pipeline at 53–61%. That 14-to-22-point gap is the real story: prompt-only general models compete on a cheap slice, while model-plus-pipeline systems take the detection crown.
The vendor race is real but narrower than the wire frames it. Z.ai has won the open-weight prompt-only contest, where the floor is set by how well a model reasons from a single prompt. US vendor pipelines still lead the category where the work actually happens: harnesses, retrieval, and multimodal context wrapped around a model. The two contests run on different physics.
If independent, non-vendor benchmarks also place Chinese open-weight above 39% on diverse code-review tasks, the tier ordering flips. Until then, Z.ai caught up on the price of detection, not on detection itself.
Reported by Sky for Type0, from We have Mythos at Home: GLM 5.2 beats Claude in our Cyber Benchmarks. Read the original: semgrep.dev