The new BioPhys-Bridge scorecard is the science community's way of saying: stop pasting biophysics papers into LLMs and treating the summary as ground truth.
The benchmark is built like a careful reviewer's checklist. Each of 500 cases carries evidence blocks with stable evidence IDs — quantitative values, units, equations, assumptions, mechanisms, the next decision the paper leads to. A model that wants to answer has to cite the right block, not paraphrase a vibe. The release spans 1,517 agent-facing tasks across six biological domains and nine physical model families, with three model families kept deliberately sparse to flag where the test is still incomplete.
The headline number: DeepSeek-V4-Flash, the leader, gets evidence citations right 0.360 times out of 1. Qwen3.7-Max lands at 0.316. GPT-4o-mini at 0.294. Every working scientist who has pasted a paper into a chat window will recognize why that ladder is low. The models are not dumb — grounding an answer in a specific evidence block, with a stable ID, in a domain where unit conversions and assumptions carry the load, is a harder task than fluent paraphrasing.
The wire will read this as "AI failed the science test." The more honest read is the opposite. BioPhys-Bridge is the field refusing to ship trust into physics-in-biology workflows on the back of marketing benchmarks. The 0.360 is the honesty floor: a deliberately hard test, three model families still marked sparse, future work explicitly to grow size and complexity. The credible response is more honest evaluation — not louder claims in either direction.
Reported by Sky for Type0, from BioPhys-Bridge: A Benchmark for Interdisciplinary Scientific Reasoning in Physics-Grounded Biological Research. Read the original: arxiv.org