Clinical AI benchmarks keep inflating on contact with real hospital timelines, and the cause is not the model: it is the data layer treating documentation time as event time. When a later diagnosis gets assigned to the moment of admission, the chart hands the model tomorrow's knowledge and the published accuracy smiles. Deploy the same model on a timeline filtered by when information was actually available, and the number collapses.
EHR2Trace's controlled experiment puts a number on the cheat. Assigning later diagnoses to the moment of admission substantially inflates measured performance; a model trained on those inflated histories then loses accuracy on availability-filtered histories, the kind a real hospital actually hands a clinician. The receipt is concrete: across three clinical datasets, the pipeline converted 846.4 million events, and detected all 28 injected faults.
The reusable move is to ask, of any clinical AI benchmark, what the timestamp actually encodes. If the answer is "when the chart was updated," the published number is measuring a model that has been handed the future, not a model that has learned to predict it. The audit trail is not a paperwork upgrade. It is the difference between a benchmark and a forecast.
Reported by Sky for Type0, from EHR2Trace: Auditable EHR Data Infrastructure for Patient World Models and Clinical Agents. Read the original: arxiv.org