Wiz is a cloud security vendor.
Wiz, a cloud-security vendor, and Microsoft published separate reports on the same Monday saying their AI bug-hunting agents now coordinate on the same workflow. Wiz's Project Atlas scored 90.9% on CyberGym, a public vulnerability-discovery benchmark, and Microsoft's MDASH scored 95.95%. Both agent pairs route per task to a different model and exchange signals at the API level rather than running side by side in silos.
Each system pairs two models on a routing decision. Wiz's Atlas pairs Anthropic's Claude Opus 4.6 with OpenAI's GPT-5.5; Wiz says Gemini is next as the company builds on its post-acquisition work with Google Cloud. Microsoft's MDASH uses MAI-Cyber-1-Flash on roughly 90% of tasks and hands the remaining 10% to GPT-5.4. MAI-Cyber-1-Flash is built on Microsoft's MAI-Thinking-1 and tuned for threat modeling, hunting, validation, and proof generation; the frontier model takes over when a task is harder.
The 90% figure comes from each vendor's own run on CyberGym, a public benchmark with disclosed methodology. The leaderboard puts a single-model frontier run on the same test in the low-to-mid 80s: GPT-5.5 Cyber at 85.6%, GPT-5.6 Sol at 83.6%, Anthropic Mythos 5 at 83.8%, Gemini 3.5 Flash Cyber in CodeMender at 83.2%. The vendor agent pairs run 5 to 12 points above the single-model frontier, and both vendors frame the gap the same way: right model for the right job, not one model for every task.
Microsoft's AI CEO, Mustafa Suleyman, says combining MAI-Cyber-1-Flash with GPT-5.4 halves customer cost versus a single frontier model. The claim is vendor-stated, not audited, and it is the kind of number a procurement officer will weight more than a researcher. Atlas, meanwhile, is internal-only at Wiz and not commercially available; MDASH is the product a customer can buy today.
Wiz published a real-world anchor for Atlas. The company said its agent pair found a critical remote code execution flaw, CVE-2026-3854, in GitHub's codebase, and characterized the resulting bug bounty as the largest GitHub has ever paid. GitHub's own statement on the bounty was not in the source material, so the size claim rests on Wiz's framing. Atlas also surfaced 200-plus previously unknown vulnerabilities in heavily audited open-source projects: gRPC, dnsmasq, Kubernetes, gVisor, the Linux kernel, and containerd.
CyberGym tests discoverability against a known catalog of bugs, and the false-positive rate of either vendor's agent pair is not disclosed in the source material. The 10% the agents miss on the benchmark is the same kind of slice that historically contains the bugs that cost the most: logic flaws, multi-step chains, and anything that needs domain context. Vendor-reported numbers on a public benchmark beat vendor-reported numbers on a private test, but they are still vendor-reported.
Wiz's Cyber Model Arena is the internal benchmarking harness the company uses to score models on threat modeling, hunting, validation, and proof generation. MDASH is the Microsoft equivalent, and it ships in Microsoft Defender today. The missing number in both disclosures is the false-positive rate; without it, the 9-in-10 figure is a catch rate, not a usefulness rate.