Everyone’s building a ‘GPT for cells.’ Relation’s bet is manufacturing the data to train it.
The business move
Relation Therapeutics, a London startup, is tackling a key bottleneck in building AI models of cells. The company plans to manufacture its own biological data to train a “foundation model of the cell,” addressing the scarcity, messiness, and inconsistency of existing data. GSK is backing this approach with up to $110 million, betting that Relation’s controlled data generation can overcome challenges holding back cellular AI.
Why it matters
AI in biology requires massive, high-quality datasets to train models capable of understanding complex cellular processes. Current biological data is fragmented and noisy, making it an unreliable basis for AI training. Relation’s strategy forces a rethink: instead of searching for better data, create it under controlled conditions that maximize consistency and scalability. This shifts the cost and effort toward data manufacturing but promises cleaner, more usable input for AI.
For drug developers, this could speed up AI-powered discovery and testing by providing models that better reflect real biological mechanisms. Investors and pharma partners gain a clearer path to AI breakthroughs without gambling on luckier data sets. The scale of GSK’s commitment highlights growing confidence in data-centric approaches to life sciences AI, pressuring competitors to improve their foundational datasets or risk falling behind.
Who gains and who gets squeezed
Relation and GSK stand to gain first-mover advantages in cellular AI and drug discovery. Pharma companies leaning on off-the-shelf biological data face rising pressure to invest in data quality or partnerships to stay competitive. AI startups depending on public or legacy datasets will find it harder to compete if Relation proves data manufacturing substantially improves model performance and reproducibility. This raises the bar for new entrants and shifts power toward players controlling data generation.
What to watch next
The critical question is whether Relation’s manufactured datasets can deliver AI models that outperform traditional approaches. Watch how GSK integrates these models into drug development pipelines and what benchmarks emerge around prediction accuracy and scalability. Also, monitor if other pharma or biotech companies follow GSK’s lead in funding data manufacturing, potentially accelerating a shift in how biological AI foundations get built.
AI Quick Briefs Editorial Desk