FAR.AI, an AI security research nonprofit, ran 1,500 automated jailbreak attempts against four leading AI models in dangerous content and cybersecurity domains, with two of them failing for under $300 of compute spent attacking.
FAR.AI, an AI security research nonprofit, launched its AI Security Leaderboard on Tuesday. The v1.0 release ranks frontier AI models on something the field has not standardized: not capability, but how cheaply model safeguards collapse under automated attack.
FAR.AI tested four frontier models across CBRNE (chemical, biological, radiological, nuclear, explosive) and cybersecurity threat categories. Each model faced 1,500 auto-generated jailbreak attempts. A "universal jailbreak" is a prompt that elicits compliant, detailed responses to more than 75% of clearly harmful questions in a domain.
Claude Fable 5 and GPT-5.6 Sol did not fail once. Gemini 3.1 Pro and Grok 4.5 produced hundreds of universal jailbreaks for under $300 of attack compute. FAR.AI frames the spread as a "hundredfold gap" between best and worst, a property of the current attack budget rather than an upper bound on model vulnerability.
The release is one day old, and the methodology is open for community feedback. FAR.AI flags what v1.0 does not yet cover: open-weight models, stronger adaptive optimization attacks like boundary-point jailbreaking, and agent hijacking.
FAR.AI's reason for the timing: regulators pulling models over cybersecurity jailbreaks, and developers holding back AI agent rollouts over adversarial-attack risk. Whether the gap is an upper bound on vulnerability remains the open question.