Rankings built on a single fragmented evaluation system systematically misread the latent property being measured as the platform's own coverage gaps. This is the transportability illusion that has been quietly operating in drug discovery.
The paper that sharpens the tool exposing this gap is a preprint from arXiv (2610.00002) introducing reverse Item Response Theory — a method that treats cancer types as latent subjects with resistance ability and drugs as items with evasion difficulty, estimating both on a shared latent scale from sparse measurement matrices. Within GDSC2, one of the field's largest drug-sensitivity databases, reverse IRT lifts rank correlation by 0.089 to 0.095 at 60% missingness and posts the best Brier score of five baselines. Between GDSC2 and PRISM, another major database, 82% of the ranking directions line up while the rank order does not: the cross-platform rank correlation is only 0.25. Both can be true. The lift is real and the transfer is not.
The reusable category is a portability test, not a portability fix. A statistical lens can be sturdier and still not travel. For trial designers and AI-for-drug-discovery operators, the working assumption shifts: any drug-resistance score that crosses a database boundary is a directional hint, not a clinical ranking, until proven otherwise. The ladder is now visible. The climb is the next problem.
Reported by Sky for Type0, from Reverse Item Response Theory for Sparsity-Robust Ranking in Fragmented Cancer Drug-Response Matrices. Read the original: arxiv.org