Xaira's 3.1 billion parameter model stopped improving on observational cell data. The fix: train on lab experiments where cells are altered on purpose, an approach few biotechs can afford.
A 3.1-billion-parameter model stopped improving on a held-out test, even as its training loss kept falling. The diagnostic came from inside Xaira Therapeutics: more parameters and more compute alone would not break the wall. The only thing that did was changing the kind of data the model trained on, moving from a giant snapshot library of cells to a much smaller collection of experiments in which scientists deliberately alter a cell and record the cascade of effects.
That pivot is the basis of X-Cell, Xaira's first virtual cell model, launched in March and detailed in a biorxiv preprint co-authored by Chief AI Scientist Bo Wang and Chief Discovery Officer Ci Chu. A virtual cell is a model that predicts how a cell will respond to a perturbation, the biological cousin of a weather model, except the input is a drug or a gene edit rather than a pressure system. X-Cell is trained on X-Atlas/Pisces, which Xaira describes as the largest genome-wide perturbation dataset to date.
Xaira's bet is that the next bottleneck in AI for biology is not parameter count or GPU time. It is information-rich, intervention-style data. That is a different bet from the one most foundation-model labs are making.
Xaira frames X-Cell as the causal successor to Wang's earlier scGPT, a virtual-cell foundation model trained on the publicly available CELLxGENE database from the Chan Zuckerberg Initiative, a single-cell corpus that has grown to more than 168 million cells and 20,000 to 30,000 gene-expression measurements per cell, as Wang described in a Decoding Bio interview. On the Latent Space episode, Wang and Chu reach for a structural-biology analogy: CELLxGENE is to single-cell biology what the Protein Data Bank is to structural biology, a foundational reference, not a training set for causal prediction.
On that episode, Chu and Wang argued the gap is between what a cell looks like and what happens when a gene is turned off or a drug is applied. Per the X-Cell preprint, a 3.1-billion-parameter model trained on a small observational dataset falls off the scaling curve, while a comparable model trained on roughly 30 times more perturbation data restores scaling. The fix was data, not architecture.
Standard single-cell datasets are passive: drop a tissue sample in a sequencer and read its gene-expression profile. Perturbation experiments are interventional. A scientist uses CRISPR to knock out a gene, applies a small molecule, or changes a growth condition, then measures the response. Each data point costs wet-lab time, reagents, and a researcher's week, which is why the world's largest single-cell snapshots have hundreds of millions of cells while the largest perturbation datasets have hundreds of thousands. The trade is between cheap observations and expensive interventions. Xaira's pitch is that interventions carry roughly 30 times more usable information per cell, enough to bend the scaling curve in a way passive observations cannot.
Latent Space's hosts estimated that the X-Atlas/Pisces data-collection and infrastructure budget runs "a few tens of millions" of dollars, with compute and headcount in the single-digit millions. Those figures are the hosts' back-of-envelope, not Xaira's. Still, the shape matches the strategy: a wet-lab bill that gates who can play, closer to the budget of a reinforcement-learning rollout than a data-rich pre-training run. Independent groups with strong ML chops but no in-house screening pipeline cannot replicate the dataset by spinning up more GPUs.
The biorxiv preprint has not been peer-reviewed, and the "largest to date" claim is Xaira's own. Independent trade-press coverage describes the launch in similar terms but does not benchmark the model against external labs. A second team that reproduces the scaling curve on a cheaper observational corpus would collapse the mechanism argument. If no one can, the cost barrier becomes the story.
AI for drug discovery has a long history of repeated hype waves, and the surviving asset from earlier rounds is mostly better target identification, not approved drugs. A model that predicts gene-expression response in a dish is not a clinical candidate. Xaira announced in July that Chu had been promoted to Chief Discovery Officer, Wang to Chief AI Scientist, and Ian McCaffery had joined as SVP of translational science, moves that point to a lab betting its own pipeline on the model rather than just licensing it. Whether virtual-cell predictions translate into compounds that work in patients is the question the preprint does not answer.
X-Cell does double duty. It is a model launch and a public, on-the-record argument that the binding constraint in AI for biology has moved from compute to data quality, and that the labs positioned for the next phase are not the same ones that won the last one. The wet-lab bill, the preprint, and the July leadership changes all point the same way. The mechanism is interesting. The replication is the open question.