AI safety guardrails were built to refuse bad prompts. They also refuse good ones, and the difference between a defender and an attacker is exactly the kind of context the guardrail cannot see.
Nvidia calls this inspection asymmetry: closed AI systems, asked to help a security team analyze an active intrusion, cannot tell a responder from an adversary, so they refuse both. Open-weight models, deployed inside the defender's own perimeter, run under controls the API never grants. That distinction became operational during a recent Hugging Face disclosure: closed-model refusals stalled the forensic analysis, while an open-weight model on Hugging Face's own infrastructure, Nvidia says, ran against more than 17,000 actions and contained the intrusion.
Nvidia's Open Secure AI alliance, formalized this week — a response Nvidia argued the incident demonstrated was necessary: that the asymmetry between open and closed AI tools has real consequences for cyber defense. Microsoft and roughly two dozen other signatories made the parallel move on the policy side with a July 24 letter to US policymakers framing openness as a safety necessity rather than a commercial preference. The notable absences of OpenAI, Google, and Anthropic from the alliance are the tell. Frontier labs that sell inference have every commercial reason to keep the asymmetry in place. The defenders are the ones who need to look inside the box.
Reported by Sky for Type0, from Nvidia pushes ahead with security alliance for AI openness. Read the original: computerweekly.com