A startup's retrospective case study recovered diagnoses a specialist lab had already flagged; that is a useful mechanism and a hard limit on what it proves.
The "AI diagnoses rare disease" headlines this week describe a result the underlying evidence does not yet support. Gamow Labs published a retrospective case study from a blinded collaboration with Pawel Stankiewicz's lab at Baylor, and what they actually showed is computational rediscovery of difficult cases the specialist workflow had already escalated. The 26 affected-infant cases in the study were not unsolved; they were unsolved at the first-line laboratory. The specialist lab solved 19. The startup's agent, George-0.1, recovered those 19 diagnoses and added two more the specialist had also flagged, leaving five unresolved.
Daniel McKinnon's segment on this week's AI:AM anchors the human stake: families with years of unresolved symptoms finally getting a name for the condition. The startup's case study also surfaces a specific filtering failure involving a distant FOXF1 enhancer, a class of regulatory element whose cell-specific role in alveolar capillary dysplasia is independently established in the peer-reviewed literature on non-coding FOXF1 enhancers. The agent is described as re-surfacing findings the clinical workflow had filtered out, not as solving cases the workflow never saw.
The evidence does not establish autonomous diagnostic solving. Two of the additional solutions Gamow reports lack independent confirmation in the published material. The company describes matching or exceeding expert performance and eliminating filtering problems; the narrow retrospective result is recovery of selected specialist-flagged diagnoses plus a handful of attributed re-surfaces. The strongest falsifier for "AI solves rare disease" would be a prospective, independent cohort where the agent finds answers the specialist workflow missed entirely. The current study is not that.
A second comparison is weaker than it reads. Gamow reports that George-0.1 correctly classified 20 of 20 healthy relatives when both systems received an affected-patient phenotype, against 6 of 20 for ChatGPT 5.5 Pro. The inputs differed: George received raw reads, the chatbot received pre-called VCF files. That is a comparison of two different pipelines, not a head-to-head model evaluation, and the deliberate phenotype mismatch for the healthy relatives further limits the read.
The benchmark side of the announcement is RareBench 0.1, a 122-case benchmark that ranks which candidate mutation is most likely the disease cause. The release reports model-dependent cost and performance results, not clinical diagnostic yield and not a total cost of care. Useful infrastructure, and a different claim.
This is where the diagnostic-AI conversation is moving from demo toward clinic, and the gap between headline and mechanism will start to matter operationally. The compute environment is part of the "why now." The same episode's AWS GPU thread reported that comparable capacity rents at roughly 3x the cost of dedicated clouds, which shapes where agentic medical-AI work runs. The Trump-Xi cooperation thread on the same episode is a qualified diplomatic signal: the source itself notes an encouraging speech is not an enforceable agreement. Both threads matter as the environment that makes retrospective reanalysis of unsolved cases economically tractable at scale. Neither is the lead.
The segment closes on a fair question: who is the slow-lane medical-AI future being built for? The families McKinnon describes have already waited years. The honest answer to whether the rediscovery mechanism scales into a prospective, end-to-end diagnostic workflow is that the current evidence does not support the marketing claim, and the fastest way to find out is to run the prospective trial the retrospective case study cannot substitute for.