🔬Causal Models Need Causal Data - Xaira’s X-Cell model for Drug Discovery (Bo Wang & Ci Chu, Chief Discovery Officer & Chief AI Scientist)
July 21, 2026
AI Summary
5 min readIn a recent episode of Latent Space, Bo Wang and Ci Chu (Chief Discovery Officer and Chief AI Scientist at Xaira Therapeutics) laid out a disciplined argument for why causal models need causal data. The core insight is straightforward but often overlooked: training a foundation model on observational, descriptive data—no matter how large—does not equip it to predict what happens when you actively perturb a cell. The team’s new model, X-Cell, is their attempt to build a virtual cell that actually generalizes to unseen contexts, and the key enabler is a massive, internally generated dataset of genome-wide perturbation experiments.
The Data Strategy: Causal Data from Perturb-Seq
The fundamental problem with most public single-cell datasets, such as CellxGene, is that they are observational. They describe what cells look like in a given state, but they do not reveal what happens when you change something—knock out a gene, activate a pathway, apply a drug. As Chu explains, “correlation data in the descriptive data set can be fit with many, many possible causal structures.” Observing that genes A, B, and C rise and fall together tells you nothing about whether A regulates B, B regulates A, or some external factor regulates all three.
Continue reading the full summary in the app — free to try.
Read Full Summary →Free • No credit card required
Never miss an episode of Latent Space: The AI Engineer Podcast
Get every new episode summarized in your inbox — free, ~5 minutes to read.
No spam. Unsubscribe anytime.
What you'll learn
- 1 (00:00) **"Wow Moment": The Model Outperforms Linear Baselines** - Bo describes the visceral impact of seeing X-Cell's predictions line up against ground truth and a linear baseline, with X-Cell clearly matching the real data.
- 2 (01:50) **Introducing Bo Wang & Ci Chu (Chu)** - The guests introduce their roles at Xaira: Bo leads Biomedical AI, Chu leads AI-enabled Discovery and the high-throughput biology group.
- 3 (03:11) **Xaira's Mission & Three AI Platforms** - The company aims to make better drugs faster using AI, built on three core platforms.
- 4 (05:58) **The Core Bottleneck: The Need for Causal Data** - The main challenge is not algorithms, but the lack of high-quality, causal biological data.
- 5 (11:14) **Defining X-Cell & "Perturbations"** - X-Cell predicts a cell's response to a genetic perturbation (e.g., turning down a gene).
- 6 (13:50) **Virtual Cell 2.0: From Equations to Data-Driven Models** - The history and evolution of the "virtual cell" concept.
- 7 (20:58) **The Critical Distinction: Causal Data vs. Observational Data** - Chu explains why a virtual cell needs perturbation data, not just descriptive profiling data.
+ Full timestamped outline available in the app
Show Notes
Bet on information
If test loss flatlines after 1.5B parameters while training loss continues to drop as you scale, that tells you that your model is limited by the amount of information in your data.
Training on a single, smallish data set exposed an information gap: the 3.1B model falls off the scaling trend. Neither parameters nor compute will improve performance past this wall. For predicting changes to gene expression, you need more information rich data.
This is what Chu and Bo’s teams have done, and here is what ~30x the information buys you:
Now we can scale with parameters and training compute! We don’t know how much this effort costed, but we can guess that data collection experiments and infrastructure was a few tens of millions, and compute + headcount + research was a few million. The budget looks like a RL rollout budget, rather than a data rich pre-training one.
We were lucky enough to have the two central figures in this story on our podcast. Taking the lead from Ci Chu and Bo Wang, Xaira Therapeutics is betting that information rich data is the key to AI-driven drug development. Chu was recently promoted to Chief Discovery Officer and Bo to Chief AI Scientist, underscoring just how strategic Xaira considers this bet.
Reverse engineering the human cell
If you had to figure out how a human cell works, what would you do? A good place to start might be by documenting what genes are expressed (e.g. what RNA is floating around) in different kinds of cells, in different circumstances.
That is CELLxGENE, a database of 168M cells built by Chan Zuckerberg Institute that maps each cell to a count of how many times 20K-30K genes were detected in that cell, plus detailed metadata about every cell. A ~4 trillion-entry matrix.
If the Protein Data Bank (PDB) unlocked structural biology models (Boltz Episode, ESM/BioHub Episode), CELLxGENE has done the same thing for Virtual Cell models. Like PDB, CELLxGENE has inspired a zoo of AI models of RNA expression; so much so that RNA expression models have become synonymous with Virtual Cell models. Bo Wang built one of the most influential, scGPT, that became the starting point for Xaira’s new model.
RNA expression ≠Virtual Cell
Models traine
More from this podcast
Latent Space: The AI Engineer Podcast →