ELMS, short for Evidence based LLM guided Monte Carlo Search, from a Georgia Tech team reuses failed protein structures as repair instructions and solved 26.7 of 30 tasks on MotifBench, a protein scaffolding benchmark, up from 16.
Today's protein-design tools generate hundreds of candidate structures and score them once; most of the expensive structural checks are then discarded. A Georgia Tech preprint turns that pattern around by keeping each failure in memory and reusing structural evaluation as repair guidance.
The framework, called ELMS (Evidence-based LLM-guided Monte Carlo Search), wraps a four-agent loop around a fixed design budget. A Critic Agent reads a failed design and names what is wrong locally, a Policy Agent picks a motif-locked operator with the parameters to fix it, and Monte Carlo tree search decides which historical failure deserves the next design effort. The structural evaluator is no longer a terminal screen; it is the instructor.
The authors (Haotian Hu, Oguzhan Gungordo, Siheng Xiong, and Faramarz Fekri) report matched-budget numbers. On the GeomMotif protocol at 100 candidates per task, ELMS reaches 86.41% success on single-motif tasks and 84.57% on paired-motif tasks, beating the strongest prior baseline by 19.3 and 21.9 percentage points. On MotifBench, it solves 26.7 of 30 tasks on average (88.89% Task Success) against 16.0 of 30 (53.33%) for the strongest baseline, under the same 100-candidate cap.
In practice, structural evaluation is no longer a filter; it is a guide. The preprint is not peer-reviewed and the benchmarks are the authors' own, so independent reproduction remains the open question before this loop becomes a default for protein-engineering pipelines.