A peer reviewed RAND study finds many dual use biology benchmark items have stopped measuring anything: the few that still do are where recent AI gains cluster.
When the latest AI model clears a bioweapon-safety test, the headline usually reads as a clean win for safety. A new RAND analysis flips that read. The test, in most of its items, has stopped being informative. The 57-page report, published August 6 by RAND's Center on AI, Security, and Technology, argues that the right question is no longer "did the model pass?" but "is this item still doing any work?"
The benchmark families in question are the dual-use biology evaluations, the closest thing the field has to a public scoreboard for AI's potential to help with bioweapon-relevant tasks. They run a model through questions that span everything from basic virology to multi-step protocol design. The problem, the RAND team writes, is that the field has been grading capability against a yardstick whose tick marks have worn out.
An item is saturated when the models being tested all clear it. That means passing no longer separates the cautious lab build from the aggressive one. The report's first finding is that this is now the majority case: many of the items in current dual-use biology benchmark suites no longer discriminate between frontier models. A score of 100 on a saturated item tells you nothing about whether the model could do the task in the wild. It only tells you the model can do the task as written.
The tool RAND uses to draw that line is item response theory, a measurement framework borrowed from educational testing. Instead of treating a benchmark as a single number, IRT models each item with its own difficulty and discrimination parameters, then places each model on a per-item curve. A saturated item is one where most evaluated models land far above the item's difficulty floor. A "discriminating" item is one where the spread across models is still wide. The report's authors, Grant Ellison, Jeffrey Lee, Barbara Del Castello, Sunishchal Dev, and Kyle Brady, used this to identify, in their words, "a discriminating frontier of tasks" that still separates today's strongest models from each other.
The second finding is the one that complicates the reassuring "AI aced the test" story. Recent model gains, the authors write, cluster on the difficult items: the ones theorized to be relevant to real-world risks, not the routine recall questions. On the easy end of the suite, model after model has long since hit the ceiling. On the hard end, each new generation moves the line. That is a different shape of progress than a flat leaderboard suggests, and it is the central reason a single pass/fail number is now misleading.
The third finding reframes what a useful comparison even looks like. Because the easy items no longer move and the hard items do, the right yardstick is between model generations: what a 2024 frontier model gets on a hard item versus what a 2026 frontier model gets on the same hard item, rather than a score against a fixed rubric. RAND's recommendation follows directly. Evaluators and the labs that run these benchmarks should publish per-item difficulty, retire or rotate saturated items, and structure future reporting so that a model's standing against its peers is at least as visible as its score against a static test.
The report is careful about what it does and does not claim. It does not say any current model can design or execute a bioweapon. It says the hard items are "theorized to be relevant to real-world risks," and the empirical case rests on benchmark scores, not on demonstrated misuse. The methodological bet is that better measurement is the precondition for any honest argument about where the danger line is. The current public scoreboard, mostly saturated and unrotated, cannot supply that argument on its own.
For the reader who follows AI safety reporting, the practical effect is small but durable. The next time a model release cites a near-perfect score on a virology or protocol-design benchmark, the question worth asking is not how high the number is but which items are still doing the discriminating. The RAND framework offers a way to ask it.