Each of the 15 flagged patients had their treatment plan changed, in what the team's Nature Medicine paper documents as a one hospital deployment, with the numbers pending independent replication.
A CT-reading artificial intelligence model from Alibaba's DAMO Academy flagged 15 liver cancers that radiologists had missed during routine reads at a Chinese hospital over a two-month period, and treatment plans changed for every patient the algorithm caught. The deployment, run at China Medical University Affiliated Shengjing Hospital and reported by Alibaba's DAMO Academy via QbitAI, is described by its developers as the first prospective real-world clinical trial of an AI designed to diagnose, not just screen for, liver cancer. The team reports in its Nature Medicine paper that the 15 missed tumors the AI surfaced averaged about one centimeter in diameter.
Yan Ke, the DAMO algorithm expert leading the work, characterized the lesions in the company-authorized reprint as "small, faint, off-position," meaning they were low-contrast against normal liver tissue or sat in anatomic locations that obscured them. In one published case, a 65-year-old bladder-cancer survivor in routine follow-up had normal liver function and tumor markers, with an initial read that flagged only calcifications. The AI red flag prompted a report revision and a chemotherapy recommendation after the algorithm identified liver metastasis.
Across more than 10,000 enhanced CT scans read during the two-month window, the team reports that AI assistance cut reading time by about 27% and lifted radiologists' sensitivity to malignancy by 11.5 percentage points. Junior radiologists using the model reached performance levels comparable to senior readers without it, a result the developers frame as a possible answer to the chronic junior-senior gap in radiology departments.
The model preserves both global liver-lesion context and local texture and boundary cues, and fuses the multiple phases of a contrast-enhanced CT scan to catch lesions that "flash" briefly during contrast wash-in and wash-out. That short window of conspicuity is the exact phase where the missed lesions the AI caught were most visible, and where unaided readers most often failed.
When the AI and the initial radiologist disagree, the workflow escalates to a senior radiologist and, if needed, to a multidisciplinary team (MDT) discussion. Every one of the 15 caught patients passed through that escalation path, which is what turned a flagged pixel into a revised treatment plan. Lesions the AI flagged but the initial radiologist did not mark were reviewed through the escalation workflow; the team's Nature Medicine paper documents the ground-truth adjudication through that clinical pathway rather than through a separate independent panel.
The model, called DAMO LiON, extends a medical-AI line that began with DAMO PANDA for pancreatic cancer on non-contrast CT, then DAMO GRAPE for gastric cancer and DAMO COCA for colorectal cancer. According to Alibaba DAMO, the earlier work produced three prior Nature Medicine papers, a National Medical Products Administration innovation-channel entry, and an FDA Breakthrough Medical Device designation for the pancreatic model. The lineage is real, but it is also vendor-supplied, and the deployment numbers should be read against that provenance.
The deployment was a single-arm, single-hospital, single-vendor study with no head-to-head control arm, no independent radiologist adjudication, and no published comparison to other commercial liver-cancer AI tools. The QbitAI-distributed report is an Alibaba-authorized reprint rather than the Nature Medicine paper itself, so the 27% reading-time figure, the 11.5 percentage-point sensitivity gain, and the 15-of-10,000 catch rate are team-reported from their Nature Medicine paper and have not yet been independently replicated. Independent replication at other hospitals, with other patient populations and other radiologist pools, is the open question the publication does not answer.
The team has flagged a multi-site expansion as the next milestone. The open question is whether the one-centimeter catches, the 11.5 percentage-point sensitivity lift, and the 27% reading-time cut hold up in centers that did not help train the model.