The collaboration sidesteps the integration problem that comes with stitching together public single cell repositories — where differences in sampling, protocols, and pipelines introduce noise — by generating proprietary perturbation data:
GSK has signed a research collaboration with Relation Therapeutics worth up to $110 million, and the headline price understates the shape of the bet. Relation's job is not to license software or hand over a finished model. The British biotech will run perturbation experiments at scale, measure how human cells respond to genetic changes and drug interventions, and feed the resulting datasets into machine-learning pipelines both companies will use to identify drug targets (artificialintelligence-news.com coverage of the GSK-Relation collaboration).
The work centers on a system Relation calls Lab-in-the-Loop. Tissue samples are profiled with single-cell transcriptomics, which reads gene activity in individual cells, and spatial transcriptomics, which keeps that readout mapped to its location in the tissue. The cells are sequenced, the targets are validated in the lab, and the results go back into machine-learning models for target identification, prioritisation, validation, and the design of the next experiment. Relation's MORGAN platform is one of those models. The loop keeps the laboratory and the model training on the same data, generated under the same protocols.
GSK and Relation have worked together before, in fibrotic diseases and osteoarthritis, and produced two functional disease datasets built on human genetics, single-cell multi-omics from human tissue, functional assays, and machine learning. The expanded collaboration extends the scope and budget of that prior work.
The deal sits against a shift in the public-data layer that most AI drug-discovery coverage does not dwell on. A 2025 review in Experimental & Molecular Medicine reports that CZ CELLxGENE, a Chan Zuckerberg-backed public single-cell repository, now provides standardised access to more than 100 million cells, and lists the Human Cell Atlas and NCBI's Gene Expression Omnibus as comparable resources. For a model trained on these corpora, raw scale is no longer the binding constraint. Integration is.
The review's authors flag the problems bulk public data carries: differences in sampling, wet-lab protocols, computational pipelines, and the noise and artefacts that come from stitching studies together. A foundation model that learns from the union inherits those gaps, and the same review notes that harmonising them is still an open research problem.
Relation's pitch is that the gap closes if you stop trying to harmonise other people's experiments and start running your own. Its perturbation work links genetic changes to disease-associated cellular characteristics and joins those results with genetic and patient-derived data. Each cycle produces a dataset that is internally consistent because the same lab, the same protocols, and the same analytical pipeline produced it.
GSK's willingness to underwrite that loop, rather than rent access to a public corpus or a third-party model, is the substantive news in the deal. The expected return is data that lets a model reason about the biology a programme is actually built on, instead of patterns averaged across heterogeneous studies that may not be measuring the same thing.
The next milestone worth watching is whether hits from the loop advance into GSK's pipeline on a faster cycle than the fibrosis and osteoarthritis programmes managed.