Clinical decision AI tools like OpenEvidence and Doximity's clinical assistant get ranked on narrow tests that say less than the headlines suggest. Here is a five question checklist for the next score.
A hospital procurement officer pulls up two pitches on a Tuesday. OpenEvidence says it beat the field on a new medical benchmark. Doximity, pitching its clinical assistant, says its numbers look different inside a real workflow. Both decks open with a single score. The officer has ten minutes and no way to tell which number actually means anything at the bedside.
That is the reader moment behind a mid-June 2026 Nature Medicine study that put clinical-decision AI tools (OpenEvidence and UpToDate Expert AI) head to head with general-purpose large language models on standardized medical questions, and the trade press reaction that followed. The paper landed and the results "rang out like a gunshot," according to STAT's health tech correspondent Katie Palmer, who used the controversy as a case study in how benchmark discourse collapses into headlines.
The practical move for a non-specialist is to set the score aside for ten seconds and ask five questions. They survive the next study, and they survive the next vendor pitch.
What was actually tested? A clinical AI benchmark is a fixed set of medical questions, usually multiple-choice or short-answer, scored against a human-curated key. The Nature Medicine comparison tested two purpose-built clinical tools and several general LLMs on the same instrument and produced a single number. Whether that number measures anything a clinician cares about is a separate question. Palmer argues, in her AI Prognosis newsletter, that a benchmark "does not mean much on its own," and that the gap between a published score and a patient is where readers get misled.
On what clinical data? Medical benchmarks live on exam-style questions, USMLE-style vignettes and board-style recall. They do not live on de-identified charts or live workflow. A tool can ace a vignette test and still stumble on a complex patient with three comorbidities. The data tells you the tool is good at the test. It does not tell you the tool is good at medicine.
Against what baseline? "We beat GPT" is not the same statement as "we beat the average clinician" or "we changed an outcome." The Nature Medicine comparison set general LLMs as the baseline, which is a meaningful bar for research but a thin one for procurement. A buyer comparing two clinical AI vendors is asking a different question than a researcher asking whether specialized models still beat generalists, and the score is being asked to do both jobs.
By whom, and reported how? Vendor-issued claims deserve a different read than peer-reviewed work. OpenEvidence has pointed to a Stanford-Harvard physician-preference study released through PR Newswire as evidence of dominance in real clinical use. That is a counter-data point worth keeping on the record, and it is also a vendor announcement, not an independent validation. ASCO AI's coverage of the Nature Medicine paper documented the same skepticism from clinical AI researchers, who questioned whether general LLMs can credibly outperform purpose-built tools on standardized exams. Both signals matter. Neither one is the whole answer.
Does the result hold up when summarized for a non-specialist? This is the gatekeeper question. If the headline is "AI beats the doctor," the underlying claim is almost certainly narrower. Palmer argues that every clinical AI study needs a headline, and the headline almost always over-summarizes the finding. A reader who can ask this question before the press release is harder to mislead.
Doximity's product page for its clinical assistant, Doximity Ask, pitches the tool as a clinical reference inside a physician workflow rather than a benchmark champion. That positioning is closer to honest than a single-score headline, and it is also the harder sell to a procurement committee. Tools that integrate into a chart and a workflow tend not to publish viral numbers. Tools that publish viral numbers tend not to publish workflow data. The procurement officer's job is to ask which kind of evidence the vendor has, and which kind the patient will eventually meet.
The Nature Medicine paper and the STAT+ analysis are useful as a case in point. Neither is a verdict on which clinical AI tool "won." Together they show how a single score gets made, summarized, sold, and pushed back against inside one news cycle. The next clinical AI headline is already in press. The five questions travel with it.
Companion coverage from STAT on the wider trust, accuracy, and safety debate around clinical chatbots lands the same point from a different door: a benchmark score is one input into a much larger clinical judgment, and the researchers who run these comparisons are usually the first to flag when a number stops meaning what the headline claims.