A new ICU framework clusters patients into five data derived physiological states from 26 routine measurements, correctly flagging the sicker patient about 86% of the time at 8 hours.
ICU monitors have a basic job: ring a bell when a number crosses a line. That works for sudden crashes. It is worse at catching the slow drift, the patient whose vitals and lab values edge in the wrong direction over hours without ever tripping a single threshold. A new peer-reviewed framework, STREAM (State Trajectory Representation and Evolution-Aware Monitoring), is built for the second case. It does not ring louder alarms. It maps where each patient is drifting, and it does it well enough to flag the difference between life and death eight hours before traditional scoring systems do.
The paper, published in PLOS Digital Health by Namvar and colleagues in 2026 (PMID 42821592), works on a layer below the bedside monitor. Every few hours, an ICU patient generates roughly two dozen routine vital signs and lab values: heart rate, blood pressure, oxygen saturation, white blood cell count, creatinine, lactate, and the rest of the standard panel. STREAM takes those 26 measurements per time point and places each patient inside a multidimensional physiological space, where similar patients cluster together and different patients sit apart. The method that does the clustering is optimal transport theory, a branch of mathematics originally developed to figure out the cheapest way to move piles of dirt into holes. Applied to ICU data, it is a way of measuring how a patient's distribution of readings compares with the distributions of other patients at the same severity level.
No one pre-labeled the states. The algorithm, run on the eICU Collaborative Research Database of 158,294 ICU stays from hundreds of US hospitals, found five reproducible physiological states on its own. Each state has a distinct clinical signature: a different mix of vitals, labs, and outcomes. STREAM then goes beyond clustering. For each new patient, the framework asks which of the five states they currently resemble, and how much of their ICU stay they end up spending inside that expected state.
That last question is where the mortality signal lives. Patients who spent less than 10% of their ICU stay inside their expected state, the "state outliers," died in the ICU at a rate of 37.6%, compared with 2.3% for "state residents," a 16-fold gap, in the development cohort. The number is not a treatment effect. It is a within-cohort stratification: the same severity scores, the same hospital systems, the same physicians, but two very different outcomes that the existing monitoring stack does not surface.
The authors then took the model to a second multicenter database, MIMIC-IV, with 84,517 ICU stays, and watched the gap reappear. State outliers in MIMIC-IV died at 33.5% in the ICU versus 3.2% for state residents, a roughly 10-fold gap, smaller than eICU but still an order of magnitude apart. The mortality-prediction model itself scored an area under the curve (AUC) of 0.863 at 8 hours and 0.903 at 72 hours in the development cohort, and 0.798 and 0.857 in the external validation, meaning it correctly identified the sicker patient roughly 86% of the time at the 8-hour mark and 90% at 72 hours. The expected calibration error was 0.002, a sign that the model's predicted probabilities matched the observed mortality rates almost exactly.
For clinicians, the second output may matter more than the score. STREAM does not just predict who is at risk. It shows which of the 26 measurements is pulling the patient toward a higher-risk state, a feature-importance analysis that points to specific labs and vitals, not a black-box "high risk" flag. That is the difference between an alarm and a diagnostic: the alarm says a number crossed a line, the diagnostic says the patient is drifting toward a state where the next number to move is creatinine.
The limits of the study travel with the result. STREAM has been validated retrospectively, on databases of patients who were already treated, and has not been tested in a prospective trial. The five data-derived states have clinical signatures, but they do not yet have a clinical meaning. The feature-importance analysis points to associations, not causes. The authors' own conclusion in the abstract stops short of any claim that the method improves outcomes, a careful phrasing that says the paper has earned the right to be evaluated, not that it has been shown to save lives.
The next round of work has to test that distinction. If a clinical team at a teaching hospital runs STREAM prospectively, watches the state-outlier alerts in real time, and sees whether acting on them changes outcomes, then the 16-fold gap stops being a stratification signal and starts being a tool. Until that study runs, the framework is a peer-reviewed capability gain for ICU monitoring, not a deployed clinical system. The open question, and the one that decides whether STREAM becomes standard of care or stays a research artifact, is whether the trajectory it sees is one a clinician can act on.