A new RAND paper argues AI is collapsing the gap between finding and weaponizing bugs in critical infrastructure code, and proposes 4 6 week sprint cycles to compress the defensive review loop.
A 24-page RAND paper out Wednesday argues the United States is letting AI find bugs in the software that runs power grids, water systems, and hospital records faster than defenders can audit it. The proposal is a tempo one. Defenders should run 4-6 week sprint cycles, not multi-year reviews, using the same AI capability that recently surfaced a 27-year-old flaw in OpenBSD's networking stack.
The paper, "Hardening Critical Infrastructure Software in an Era of Rapid AI Advancement," was published Aug. 27 by RAND's Center on AI, Security, and Technology. Its authors are Gopal P. Sarma, Rachel Steratore, Sunny D. Bhatt, Greg McKelvey Jr., and Michael Jacob. The work was sponsored by the Hewlett Foundation and draws on structured literature review, LLM-assisted research, and conversations with more than two dozen experts across academia, industry, and government.
The framing is dual-use. AI is compressing the offensive timeline against aging critical-infrastructure code, the paper argues, but the same capability can be turned to defense if policymakers organize to do so. The bottleneck is no longer tooling. It is the cadence of human review.
That is what the 4-6 week number is buying. Each convening would pull together AI-assisted code analysis specialists, infrastructure operators, and high-assurance software engineers to look at one slice of critical-infrastructure code, return with concrete guidance, and repeat. RAND calls the model a "rapid technical convening" and wants it to run repeatedly rather than as a one-shot. The first cycles would review findings from Glasswing and similar programs before moving into operator-led assessments.
The proposal lands in a market that is already moving. Anthropic announced Project Glasswing on April 7, 2026, with twelve founding partners including AWS, Apple, Broadcom, Cisco, CrowdStrike, Google, JPMorganChase, Linux Foundation, Microsoft, NVIDIA, and Palo Alto Networks. The program's flagship model, Mythos Preview, has surfaced a 27-year-old signed-integer overflow in OpenBSD's TCP SACK implementation that leads to a null-pointer dereference, a 16-year-old type mismatch in FFmpeg's H.264 decoder that produces an out-of-bounds heap write, and Linux kernel privilege-escalation chains. A Cloud Security Alliance whitepaper put Glasswing's backing at $100 million in credits plus $4 million in direct funding.
The benchmark jump is sharper than the model lineup suggests. On CyberGym, a standard AI vulnerability-discovery test, Mythos Preview scored 83.1% versus 66.6% for Anthropic's own Opus 4.6. RAND's paper flags that as more than incremental. The qualitative difference is that the higher-scoring model surfaces real, exploitable bugs in real code, not just synthetic test cases. That is the kind of capability the convenings are meant to bring to bear on operator-owned systems.
The Defense Advanced Research Projects Agency ran a smaller version of this experiment. The AI Cyber Challenge ran from August 2025 to March 2026, and according to Cybersecurity Dive, the three winning teams found 83 vulnerabilities across more than 30 projects including Android, Linux, SQLite, and Redis. Atlanta, Trail of Bits, and Theori/42-b3yond-6ug shared roughly $830,000 of a $1.4 million bonus pool, and the organizers noted that logic bugs, the kind that are hardest for static analyzers, were particularly amenable to AI discovery. Trail of Bits has since partnered with HHS/ARPA-H to apply the same toolchain to medical-device firmware, which is one of the more direct critical-infrastructure adoption paths so far.
The RAND proposal is a way to take the AIxCC pattern out of the competition setting. The convening is meant to produce concrete, prioritized findings that an operator can hand to their engineering and procurement teams, not a leaderboard. Each cycle is short enough that the same specialists can rotate through multiple infrastructure sectors. That is the bet. AI has compressed the offensive review loop. Defenders need to compress the defensive one.
The catch is downstream. A sprint can surface a 27-year-old OpenBSD bug, but the next step is patching production systems where downtime is a physical risk: a substation controller that cannot be rebooted mid-shift, a hospital infusion pump on a regulated update cadence, a water treatment SCADA system that is air-gapped for a reason. Procurement, legal review, and reliability validation run on different clocks than AI does. RAND's 4-6 week number does not fix those. It only puts the same AI capability into a tempo defenders can actually use, and assumes the rest of the stack catches up.
The next convening the paper envisions has no host yet. RAND's bet is that one of the program partners, or a sector coordinating body like the Electricity Information Sharing and Analysis Center, picks it up before the offensive tempo pulls further ahead.