OpenEvidence and UpToDate Expert AI were approved as medical devices. A Nature Medicine benchmark from NYU Langone found GPT 5.2, Gemini 3.1, and Claude Opus 4.6 beat them on the study's own tests — the comparison the FDA never required.
The FDA cleared OpenEvidence and UpToDate Expert AI as medical devices. It never tested whether they were the best tool for the question a doctor is actually asking.
A Nature Medicine paper from NYU Langone Health published 12 June 2026 did. Across three sets of clinical questions, the general-purpose models from OpenAI, Google, and Anthropic, none of them FDA-cleared but all of them already in doctors' pockets via consumer apps, outperformed the regulated clinical tools.
The result has rattled the medical-AI community. A Kaiser Permanente vice president for AI and emerging technologies wrote on LinkedIn that she had "never seen a single paper trigger the kind of reactions this one has in the health AI community." Many of those reactions read the result as a clean win for general-purpose models. The paper's authors are more measured, and so should the conclusion be.
The benchmark was designed to compare the tools on the same questions, not to crown a winner. Alyakin, Oermann, and colleagues at NYU Langone built a three-stage evaluation: 500 MedQA knowledge items, 500 HealthBench clinician-alignment items, and a real clinical queries (RCQ) benchmark built from 100 de-identified questions submitted by physicians to a HIPAA-compliant GPT instance at the hospital. Twelve U.S. clinicians produced 1,800 blinded model-question annotations.
In all three stages, the three frontier general-purpose models (OpenAI's GPT-5.2,) beat OpenEvidence and Wolters Kluwer's UpToDate Expert AI. The clinical tools performed about as well as Google Search's auto-enabled AI Overview on the RCQ benchmark. That is the part of the result the online reaction skipped past.
The structural problem is what the FDA was asked to evaluate. The clinical AI tools moved through a regulatory pathway built on the 4 December 2024 final guidance on Predetermined Change Control Plans for AI-enabled device software. That pathway asks whether a sponsor's device meets the sponsor's own performance specifications, not whether the device beats every general-purpose alternative a doctor might use instead. The FDA's AI-Enabled Medical Devices List catalogs cleared tools by submission number and decision date. It does not document comparative clinical performance against general-purpose models.
The FDA's clearance answered one question. The NYU Langone benchmark asked a different one. The two questions do not overlap the way the marketing copy suggests.
STAT+'s 29 July 2026 analysis framed the result as a moment for clinicians to think about which tools they trust at the point of care. Becker's Hospital Review reported the headline finding as a "ChatGPT, Gemini, Claude beat clinical AI tools" result. A ClinicalTrialVanguard opinion piece read the gap as a regulatory validation failure.
The counterargument is real. OpenEvidence, UpToDate, and Doximity are not pitching a better chatbot. They are selling integration: an audit trail, a citation footnote attached to every answer, an interface that lives inside the electronic health record, and a vendor contract that a hospital's risk department can enforce. None of that shows up in a multiple-choice benchmark or even in a 100-question RCQ test. A clinical AI tool that scores lower on accuracy can still be the right tool for a health system that has to defend its decisions to a regulator, a payer, or a malpractice lawyer.
The benchmark also has limits the authors flag. The RCQ corpus came from a single hospital's queries, scored by 12 clinicians in one institution. Replication, not victory laps, is the next step.
Hundreds of thousands of U.S. clinicians already use clinical LLMs as a hedge against hallucination-prone generalist models. The Nature Medicine result does not mean the hedge is worthless. It means the hedge was never measured against the model it was supposed to keep doctors away from. The FDA cleared the clinical tools without that comparison, and the comparison is what the next regulatory guidance will have to add.