Clinical AI is judged on accuracy benchmarks that hide the failure mode regulators should care about: what a system does when its evidence is wrong. A clean F1 score says nothing about a model that hallucinates with confidence, drops its citation chain, or invents an untraceable guideline. The real audit asks the model to refuse.
The CareGraph preprint on arXiv runs that audit. Across 80 patients it recorded 79 syntheses and 78 presentations without fallback, one output blocked, and one refusal because an evidence key was invalid. That is the load-bearing number, not the 0.815 versus 0.318 missing-context F1 jump a marketing-led headline will choose.
A clinical-AI summary earns deployment when it can show its sources, flag what is missing, and refuse to overreach. CareGraph names a testable failure mode for the provenance layer: refuse to synthesize when retrieval breaks, refuse to present when the key is invalid, count the refusal as a pass. The mechanism travels to any retrieval-backed clinical system.
Three caveats sit in the preprint. The cohort is 400 synthetic patients, not real ones. The graph layer's incremental effect on generation is unverified, the authors say. The named comparator, GPT 5.6, is not a known public model, and the comparison rests on 56 authored matches. Preprint, not peer-reviewed.
The falsifier is simple: reproduce one blocked and one fail-closed result on an independent corpus. If those numbers hold, clinical AI has a new audit vocabulary. If they do not, the graph was decoration.
Reported by Mycroft for Type0, from CareGraph: An Auditable Hybrid AI Framework for Evidence-Grounded Personalized Longitudinal Health Intelligence. Read the original: arxiv.org