A Florida pastor is suing OpenAI, saying ChatGPT delayed his pulmonary embolism care. The same day, the company's health lead tempered its 'better than clinician level' claim.
Ashley Alexander, OpenAI's vice president of health product, said in a prelaunch briefing this week that the company's models are "now capable of reasoning at levels that are better than clinician level." One follow-up question later, the company's own health lead was tempering that claim on the record.
Alexander made the statement as the company prepared to open ChatGPT Health to every logged-in US adult on web and iOS. The Verge first reported the rollout and the quote, then asked Karan Singhal what to make of it. Singhal said he would "temper" the framing, pointing to "individual studies that have been pointing in that direction" rather than a settled clinical claim.
The rollout runs on GPT-5.6 Sol, the model OpenAI introduced as its flagship health model. ChatGPT Health now lives inside the main chat window rather than a separate tab. OpenAI's own disclaimer reads that ChatGPT "supports, not replaces, professional care." That line sits in the same launch announcement as Alexander's "better than clinician level" claim, and the company has not retracted either one.
A Florida pastor filed suit against OpenAI on Wednesday alleging the chatbot gave him "extremely dangerous medical recommendations" that delayed treatment for a pulmonary embolism, a blood clot in the lungs that turns life-threatening when treatment stalls. The lawsuit is active litigation; the allegations are what the pastor's lawyers have pleaded, not adjudicated facts. The case lands on the same day as the public rollout and inside the same press cycle as the claim Singhal walked back.
The evidence OpenAI leans on for clinician-level performance is partly its own. HealthBench Professional, the benchmark the company cites, is described in a paper hosted on both the OpenAI CDN and arXiv. arXiv preprints have not been peer-reviewed, and HealthBench Professional is a company-built evaluation: OpenAI designed the cases, OpenAI graded the responses, and the doctor-comparison set is narrow. Singhal also pointed to a Harvard and Stanford study covered by Harvard Magazine and Fortune, in which OpenAI's earlier o1-preview model outperformed attending physicians on a small set of ER diagnostic vignettes. Those vignettes are written scenarios, not real patients, and are narrow proxies for clinical practice rather than a substitute for it.
The pattern across OpenAI's clinicians post and its ChatGPT Health launch page is consistent: a capability claim at the top, with the proof being a benchmark the company built and a small academic study on simulated cases. Singhal's walk-back is unusual for being on the record and for coming inside the same press cycle as Alexander's claim.
The portable test for the next AI health launch: when a company says "clinician level," the next sentence to listen for is what its own people will defend on the record. If the defense is "individual studies pointing in that direction" and a company-built benchmark, the claim is the marketing and the tempering is the substance. The Florida lawsuit, whether or not it survives, is the reason that gap matters at all.