A systematic review of 15 studies finds one digital pathology platform already validated across 12 centers, but most AI models still lack the testing needed for clinic.
For the thousands of patients diagnosed each year with non-muscle-invasive bladder cancer (NMIBC), the early-stage form that has not yet reached the muscle wall, the standard treatment is an odd inheritance. The drug is BCG, or bacillus Calmette-Guérin, a weakened tuberculosis vaccine that has been poured directly into the bladder for decades. It works for most patients. For a large minority, it never will, and they only learn that after months of treatment and follow-up.
A new systematic review of 15 artificial-intelligence models built to predict who will and will not respond to BCG finds that one tool has cleared the harder test: external validation. That tool is CHAI, a digital pathology platform that reads standard tissue slides and has been tested across 12 international medical centers. The other 14 models are still research-grade, including five that the review's own risk-of-bias assessment rated as high concern.
The review, published in BJU International and registered with PROSPERO before the search began, combed PubMed, EMBASE, and Web of Science through April 2026. It covered more than 24,900 patients across the 15 studies, split between seven "perceptual" models that read histology or imaging directly and eight "integrative" models that combine clinical, pathological, and molecular data.
CHAI's hazard ratios are the strongest signal in the field. Patients whose tissue the platform flagged as high-risk were about twice as likely (HR 2.08) to see their cancer recur at a high grade, nearly four times as likely (HR 3.87) to progress to a more dangerous stage, and roughly two and a third times as likely (HR 2.31) to be classified as BCG-unresponsive as patients the platform scored as low-risk. In a cohort where some patients received gemcitabine or docetaxel, chemotherapy alternatives that are filling the gap left by a multi-year BCG shortage, CHAI's score also distinguished patients who benefited specifically from BCG rather than from the alternatives (interaction P = 0.029). It tells clinicians which patients are likely to benefit from BCG, sparing the rest months of ineffective treatment.
The other 14 models in the review posted smaller gains. Two integrative tools, PROGRxN-BCa and a DeepSurv-based model, produced concordance scores of 0.79 and 0.881, beating the standard risk calculators that urologists use today, but only by five to ten percentage points. Concordance is a measure of how well a model ranks patients by risk. None of these results came from a prospective trial.
External validation is the difference between a model that works in the lab and one that holds up in a different hospital, on a different scanner, with a different patient population. Only 40 percent of the 15 studies had performed any external validation at all, and none had run a prospective trial. Five studies, on the review's PROBAST assessment, were at high risk of bias, a flag the authors raised rather than buried.
Merck, the main supplier of Tice BCG, has acknowledged years of rationing, and the Bladder Cancer Advocacy Network has tracked the downstream effects on patient care for most of the last half-decade. The CIDRAP news site has reported that thousands of US bladder cancer patients have received substandard care as a result. With gemcitabine and docetaxel moving into practice, knowing in advance which patients truly need BCG and which can be steered elsewhere has become a clinical question with real consequences.
The review's authors do not claim CHAI is ready to deploy. They call instead for prospective trials, the kind of study that puts the model in front of real patients at multiple centers and measures whether using it changes outcomes, not just predictions.
That is the next bar. CHAI is the only one of 15 models over the first one.