RAND's first public test of LLM agents against the biosecurity gate on custom DNA orders found emerging, inconsistent evasion, with the most advanced commercial AI systems — whose internal workings aren't public — outside the experiment.
RAND's Center on AI, Security, and Technology asked a specific question this month: can a large language model agent act like a less-expert version of a trained molecular biologist, redesigning a short DNA order well enough to slip past the screening that custom synthesis vendors run on every incoming sequence? The answer, in a 48-page peer-reviewed report published August 11, 2026, is a hedged yes. Agents can sometimes route around the gate, sometimes fail, and sometimes refuse to try.
Nucleic acid synthesis screening, often called NASS or DNA-synthesis screening, is the check vendors of custom DNA and RNA run on every sequence a customer orders. If the order matches a known pathogen genome or a fragment of one, the vendor blocks it. The system was built for a world of trained bench scientists filing orders by hand. It was not built for an LLM agent that can iterate on a redesign until the sequence still encodes the right protein but no longer trips the homology filter.
RAND's test isolated that bottleneck. Researchers had LLM agents attempt to redesign peptides and proteins so the encoding DNA would preserve function while evading common screening. Agents succeeded on some tasks and failed on others. They also hit a hard ceiling: model safeguards prevented the team from probing closed-weight frontier systems directly, the proprietary models whose internal safety filters block certain biosecurity-relevant outputs. The result is a partial read of the threat surface rather than a complete one, and the report flags the gap explicitly. A June 2026 companion report from the same research group found that LLM agents can already perform initial interactions with biological tools, which is the lower rung of the capability stack the August report then escalates.
Independent corroboration came from a separate benchmark, ABC-Bench, accepted at ICML 2026 and posted to arXiv in June 2026. Across Fragment Design, Liquid Handling Robot, and Screening Evasion tasks, tested LLM agents outperformed the median expert human baseliner. On the Screening Evasion task specifically, however, Claude Opus 4.6, Claude Sonnet 4.6, and GPT-5.4 refused every sample. The agents that did try were the open-weight and older ones, the same population RAND had to lean on because of the safeguard ceiling.
The mechanism behind successful evasion is design-layer, not end-to-end. A 2025 NCBI Bookshelf synthesis on AI-enabled biosecurity describes how tools like ProteinMPNN can redesign a protein to keep its structure while swapping amino acids, which changes the DNA that encodes it. Homology-based screening compares a customer's DNA to known pathogen sequences; change the codons, and the match softens. The protein still folds. The pathogen resemblance drops below the flag threshold. As the same NCBI synthesis notes, current AI tools can facilitate the design of simple biomolecules such as toxins; they do not enable de novo design of a self-replicating organism. The physical design-build-test-learn loop, the actual wet-lab work of synthesizing and testing a candidate, is still where the bottleneck sits.
The threat picture changes depending on which layer AI is helping with. The 2024 RAND Red Team study, OpenAI's 2024 evaluation, and Anthropic's 2024 work all reached milder conclusions: no statistically significant difference in attack plan viability, "at most a mild uplift" for GPT-4, and Claude 3 uplift only for novices in "certain parts" of acquisition. RAND's August 2026 report does not contradict those findings; it narrows the scope to a specific gate and finds emerging capability there, while leaving the larger end-to-end question open.
The US framework for nucleic acid synthesis screening is mid-revision under Executive Order 14292, and the Biosecurity Handbook describes AI as a risk amplifier rather than a risk creator. Vendors that run the screening, model deployers whose safeguards blocked RAND from probing closed-weight systems, and policymakers redrawing the rulebook now have an empirical, if partial, picture of how the gate holds under AI-assisted pressure.
The August report narrows the question to a specific gate, finds emerging capability at the design layer, and leaves the closed-weight frontier and the end-to-end question open. The E.O. 14292 revision is where those open questions turn into rules.