The lab is the first to use its 2023 Preparedness Framework's top tier as a public flag, applying tighter internal safeguards while it benchmarks further.
OpenAI says an unreleased model it has been evaluating internally may have crossed the company's own "Critical" cyber threshold, and is telling the public before deployment. The voluntary flag is the first real-world test of the Preparedness Framework OpenAI published in 2023, and the closest thing the AI industry has to a named bar for handing a capability jump to the world.
Under that framework, "Critical" cyber means an AI can identify and develop functional zero-day exploits in hardened real-world systems without human help, or design and execute end-to-end novel attack strategies from a high-level goal. Zero-day exploits are previously unknown software flaws attackers can use before a vendor patches them. The prior model rated at this tier, GPT-5.6-Sol, was assessed at High rather than Critical under the same scale.
OpenAI wrote that internal evaluations of an upcoming model showed "significant advancements in agentic coding and cybersecurity" and that it "cannot rule out" the Critical tier. The lab says the same model was not tied to a recent Hugging Face exploitation incident.
The disclosed safeguards are internal: isolated testing, restricted network and tool access, stronger model-weight protections, more monitoring, sandboxed execution, and a pause on some activities. No external audit is attached, and both the threshold and the assessment remain OpenAI-defined.