The most important paper in AI virtual cells is a negative result with a blunt title: deep-learning-based gene perturbation effect prediction does not yet outperform simple linear baselines.
The setup was designed to give the models every advantage. Norman 2019 CRISPR-activation data, 100 single genes and 124 gene pairs in K562 cells. Each model was fine-tuned on all the singles plus 62 of the doubles, then tested on the 62 held-out doubles across five random splits — the favourable regime, not a hostile one. Error was measured as L2 distance over the 1,000 most highly expressed genes.
Seven models entered: five single-cell foundation models (scGPT, scFoundation, scBERT, Geneformer, UCE) and two task-specific deep models (GEARS, CPA). Against them sat two deliberately trivial baselines — “no change” (predict the control profile) and “additive” (sum the single-perturbation log fold changes).
The published result, verbatim: “All models had a prediction error substantially higher than the additive baseline… None of the models was better than the ‘no change’ baseline.”
There is a serious rebuttal, and it deserves to be stated as strongly as the negative result. Miller et al. argue in Nature Biotechnology that the negative reports “largely stem from limitations of benchmarking metrics, not from the models themselves.” They introduce a positive-control baseline and a calibration measure, identify two failure modes — control bias and signal dilution — and report that across 14 datasets and up to 18 metrics, “under well-calibrated metrics, deep learning models outperform mean, control, and linear baselines, and in some cases even surpass the additive baseline.”
That rebuttal does not overturn the experiments. It disputes their measurement. And the fact that the field’s central question has become “is your metric calibrated?” rather than “is your model bigger?” tells you where this technology actually stands.
What a virtual cell is supposed to be
The canonical framing is Bunne et al. in Cell (December 2024): an AI virtual cell is “a comprehensive AI framework composed of several interconnected foundation models that represent dynamic biological systems at increasingly complex levels of organization—from molecules to cells, tissues, and beyond.” It has two parts: universal representations — a shared embedding space spanning scales and species — and virtual instruments, networks that decode or manipulate those representations to run in silico experiments.
A newer formulation pushes further. Xing & Song, in Cell (September 2026), define the virtual cell as an action-conditioned generative world model that “simulates biological possibilities of a cell under any natural or artificial interventions” — explicitly contrasting that with predictive foundation models that perform “specific tasks, such as gene-expression perturbation prediction.”
That definitional drift is worth holding onto, because it is where most of the confusion in this field originates. A task-specific predictor and a world model carry completely different evidentiary burdens. A great deal of marketed capability uses world-model language while delivering a task predictor — and the benchmarks measure the task predictor.
┌─ MEASUREMENT ───────────────────────────────────┐
│ scRNA-seq · scATAC · proteomics · spatial │
│ imaging · perturbation screens │
└────────────────────────┬────────────────────────┘
▼
┌─ REPRESENTATION — works today ──────────────────┐
│ shared latent cell space │
│ Geneformer · scGPT · UCE · STATE │
└────────────────────────┬────────────────────────┘
▼
┌─ TRANSITION — unproven ─────────────────────────┐
│ f(state, perturbation, context, time) → state' │
│ <-- the additive baseline wins here │
└────────────────────────┬────────────────────────┘
▼
┌─ VALIDATION ────────────────────────────────────┐
│ held-out perturbations · unseen cell types │
│ organoids · independent cohorts │
└────────────────────────┬────────────────────────┘
The representation layer is genuinely working
This part of the stack is real, funded, and largely peer-reviewed. The models that matter, with verified figures:
| Model | Scale | Status |
|---|---|---|
| Geneformer | ~30M transcriptomes, rank-value encoding | Nature 2023 |
| scGPT | 33M+ cells | Nat. Methods 2024 |
| scFoundation | >50M profiles, 100M params, read-depth-aware | Nat. Methods 2024 |
| UCE | 36M cells, >1,000 cell types, 8 species, no fine-tuning needed for new cells | Nature 656:183–191, Jul 2026 |
| TranscriptFormer | up to 112M cells, 12 species, 1.53 Gyr of evolution | Science |
| STATE | 167M observational + >100M perturbed cells across ~70 contexts | Cell, Aug 2026 |
| AlphaGenome | 1 Mb sequence input, 25 of 26 variant-effect evaluations matched or beaten | Nature, Jan 2026 |
| C2S-Scale | 27B params, cells rendered as text for an LLM | Preprint + blog only |
| X-Cell | 4.9B params, 25.6M perturbed cells, 7 CRISPRi screens | Preprint only |
Three details in that table are worth more than the headline numbers.
Geneformer’s rank-value encoding is routinely misdescribed. Genes are not ranked by relative expression within the cell. They are ranked by expression in that cell scaled by expression across the entire ~30M-cell corpus. That corpus-wide scaling is the mechanism — it suppresses housekeeping genes and promotes transcription factors. Strip the scaling out of the description and the model looks arbitrary when it is deliberate.
STATE’s cell counts resolve cleanly once you separate the modules. 167M is the observational cell count for the embedding (SE) module — the exact figure in the paper. “Nearly 170M” is Arc’s own rounding of that same number. The 267M figure circulating in summaries is 167M + 100M, a derived total that is not a headline figure in the paper at all. The perturbational module used over 100 million perturbed cells across ~70 contexts.
STATE also carries the most interesting architectural claim in the field. Because sequencing destroys the cell, a given cell’s pre-perturbation state is unobservable — the paper defines X₀ as “drawn from a basal cell distribution 𝒟_b.” So the model is trained to predict the perturbation effect against the whole basal population distribution, not against an individual cell. That is a causal framing, and it is more careful than most.
Underneath all of it sits data. Arc’s Virtual Cell Atlas exceeds 600 million cells across observational and perturbational data. Tahoe-100M holds 100M transcriptomes from 50 cancer cell lines across 1,100 drug-dose conditions — worth stating precisely, because those 100M transcriptomes come from roughly 60,000 perturbation experiments, not 100M distinct perturbations. A further 120M+ cells and 225,000 perturbation interactions from Tahoe, Arc and Biohub were announced in January 2026 for open release; as of October they have not shipped.
The evaluation problem
This is the part of the field that most summaries omit, and it is the part that determines whether any of the rest is usable.
The competition result. Arc’s Virtual Cell Challenge 2025 used ~300,000 single-cell profiles from 300 CRISPRi perturbations in H1 human embryonic stem cells. Arc’s own wrap-up states that models were “not yet consistently outperforming naive baselines across all metrics,” that winning approaches “combined deep learning with classical statistical features” — and that “almost all models performed worse than baseline on MAE,” so the top performers “strategically focused their efforts on PDS and DES.”
Read that last clause again. The organiser is describing metric gaming, plainly, about its own leaderboard.
A citation error worth correcting. The “2025 Nature Methods evaluation of single-cell foundation models” is Genome Biology 26:101 (Kedzierska et al.). Nature Methods ran a news highlight of it. The study itself evaluated two models — Geneformer and scGPT — narrowed down from ~12 candidates because code and weights were unavailable for the rest. That narrowing is itself a finding: what can be benchmarked is filtered by what publishers release, which means the published negative literature systematically over-samples the models that ship.
Split design changes the answer. Mao et al. benchmarked 11 methods across ~7 datasets from 6 studies, building unseen-cell splits in a foundation-model embedding space by cell-state similarity rather than at random. Performance “decreases across nearly all methods under the scFM-based split relative to the random split,” with the authors’ conclusion stated flatly: “random splitting can make models look more robust than they actually are.”
The intervention finding is the one to build on. Ahlmann-Eltze and colleagues did not stop at the negative result. They hypothesised the failure was partly because the pretraining data is observational, and demonstrated that a simple linear model pretrained on a perturbation dataset reliably outperformed every other model — including across cell lines. Their conclusion: atlas pretraining gave “only a small benefit over random embeddings, but pretraining on perturbation data increased predictive performance.”
That single sentence reframes the entire data strategy. More observational cells does not fix this. Interventional data does.
And the clean read is imminent. Arc’s 2026 Challenge is zero-shot: no challenge-specific training set, six cell lines never seen perturbed, inputs limited to non-targeting-guide profiles plus a list of gene IDs to knock down, ground truth withheld. Final test data releases 22 October 2026, submissions are due 5 November, and winners are announced in mid-to-late November. If the field has moved past its 2025 baselines, that is where it will show.
What has actually been validated
| Result | Validation status |
|---|---|
| UNAGI predicted nifedipine as anti-fibrotic in IPF; confirmed in human precision-cut lung slices with a fibrotic cocktail | Peer-reviewed (Nat. Biomed. Eng. 2025), but same group |
| C2S-Scale + silmitasertib: ~50% increase in antigen presentation with low-dose interferon in two neuroendocrine models unseen in training | Preprint + company blog; same-group validation |
| Geneformer in silico knockouts across 109 cell types → GSTP1 as an NSCLC immunotherapy-resistance driver | Peer-reviewed, single lab |
| insitro ALS/TDP-43 → 3 targets with a $25M BMS milestone; shared latent dysfunction across C9orf72/VCP/TDP-43 models | Company-reported; targets undisclosed |
| ProteinTalks: >38M temporal protein measurements; efficacy, synergy, resistance | Peer-reviewed (Nature, Sep 2026) |
| VirTues: spatial proteomics; predicted anti-PD-L1 response and stratified disease-free survival in an independent cohort | Peer-reviewed — the only entry with independent-cohort validation |
| AetherCell: teriflunomide for dry eye, dabigatran for ulcerative colitis | Unverified preprint; the “first virtual-cell model with in vivo validation” claim is not substantiated by any source |
Two numeric corrections in circulation: MolPhenix’s reported improvement over prior phenotypic models is 8.1×, not 10×; and the widely quoted “136 molecules in under 12 months” describes REC-617 only — Recursion’s REC-1245 took 18 months and 204 novel compounds.
Note the shape of the table. Every result except one is same-group validated, company-reported, or preprint-asserted. The honest summary is not that the work is bad — it is that the validation architecture is still internal to the groups making the claims.
The regulatory position — and it is not what the summaries say
This is where circulating commentary diverges most sharply from the actual instruments, and it matters because it is the difference between a plan and a wish.
| Instrument | What it actually is |
|---|---|
| FDA Modernization Act 2.0 (enacted 29 Dec 2022, §3209 of FDORA) | Replaced “preclinical tests (including tests on animals)” with “nonclinical tests” in FD&C Act §505(i)(1)(A), made a parallel edit to the biosimilar pathway, and added §505(z) defining “nonclinical test” as in vitro, in silico, in chemico, or nonhuman in vivo. Permissive, not prohibitive — it did not ban animal testing and did not amend 21 CFR 312.23(a)(8), which still reads “laboratory animals or in vitro” |
| NAMs draft guidance (CDER, March 2026; 91 FR 13313, Docket FDA-2025-D-6131) | A non-binding validation framework for in vitro and in silico new approach methodologies. It does not make NAMs “primary evidence” — it encourages them within weight-of-evidence, and notes a NAM need not even be validated to be considered. Safety/toxicology scope |
| “Nonclinical Testing Terminology” Direct Final Rule (22 Sep 2026; 91 FR 59988; effective 4 Feb 2027) | A terminology rule: replaces “animal” with “nonclinical” across 21 CFR 312, 314, 315, 361, 601. FDA states it “does not change evidentiary standards or impose new costs” |
| AI credibility draft guidance (Jan 2025, Docket FDA-2024-D-4689) | The seven-step risk-based credibility framework: define the question → define context of use → assess model risk → credibility plan → execute → document deviations → determine adequacy. Still draft and non-binding; covers safety, effectiveness and quality |
| ICH M15 model-informed drug development | General principles for model-derived evidence; harmonised, non-binding |
| FDA + EMA guiding principles of good AI practice (14 Jan 2026) | Ten principles, intended to “underpin future AI guidance”; non-binding |
Two corrections that follow.
First, scope. FDMA 2.0, the March 2026 NAMs draft and the September 2026 rule are all nonclinical-safety instruments. The September rule expressly excludes the Animal Rule and orphan-drug provisions. Only the AI credibility guidance and ICH M15 reach efficacy and disease modelling. A cell-based or in silico model replacing a toxicology study is a different proposition from one replacing a clinical efficacy claim, and only the first has any instrument behind it today.
Second, and more simply: no FDA or EMA instrument names digital twins, virtual control arms, or a virtual cell, and none creates a dedicated evidentiary pathway for one. The nearest general vehicles are ICH M15 for model-derived evidence and the January 2025 AI credibility draft for AI-derived evidence.
The practical consequence is not that virtual cells are unusable in a submission. It is that each one gets assessed piecewise against generic model-evidence frameworks, per context of use — which means “it’s a foundation model” is not evidence, and there is no AI-specific fast lane to design around.
If you are building on this
- Choose your baseline before you choose your model. A no-change predictor and an additive-sum predictor are the floor. If a result does not report both, the number is not interpretable.
- Report split design explicitly. Random splits flatter; unseen-cell splits in embedding space are the honest test. Most published gains shrink under the second.
- Prioritise interventional data over observational scale. The single strongest empirical finding here is that a linear model pretrained on perturbation data beat every foundation model. Fine-tune on perturbations.
- Treat calibration as a deliverable, not a diagnostic. The metric dispute is unresolved, and it will not be settled by the field — it will be settled by whichever evaluation protocol survives independent replication.
- Label provenance on every claim. Peer-reviewed; peer-reviewed but single-lab; preprint; company-reported. Every flashy application result in this space sits in the last two buckets, and the only independently validated one is a spatial-proteomics model.
- Scope the regulatory claim before the scientific one. Know whether the decision your model informs is a safety decision (there are instruments) or an efficacy claim (there is essentially nothing specific), and validate to that context of use as the January 2025 framework requires.
The bottom line
AI virtual cells have a real and working representation layer, a well-funded data layer, and an unproven transition layer. The field knows this — its own benchmarks say so, its own competition organiser said so, and its most important paper is titled as a negative result.
The rebuttal is legitimate and unresolved, but notice what it concedes: that the disagreement is about measurement, not capability. That is the actual state of the art — not “AI models the cell,” but “we are still arguing about how to score whether AI models the cell.”
The 2026 Virtual Cell Challenge is zero-shot, on six cell lines never seen perturbed, with results in November. That is the number to watch. Everything being marketed today should be read as a hypothesis generator until it lands.
Related: 117 AI-Built Drug Programs, Zero Approvals · The Generalization Wall: What Molecular AI Actually Delivers · Knowledge Graphs Without Hallucinated Edges
Sources: Ahlmann-Eltze, Huber & Anders, “Deep-learning-based gene perturbation effect prediction does not yet outperform simple linear baselines,” Nature Methods 22:1657–1661 (2025) · Kedzierska et al., Genome Biology 26:101 (2025) · Miller et al., Nature Biotechnology (Oct 2026) · Mao et al., arXiv:2604.27646 · Wei et al., Nature Methods (2026) · Bunne et al., Cell 187(25):7045–7063 · Xing & Song, Cell 189(19):5831–5844 · Theodoris et al., Nature (2023) · Cui et al., Nature Methods (2024) · Hao et al., Nature Methods (2024) · Rosen et al., Nature 656:183–191 · TranscriptFormer, Science · STATE, Cell (2026) · CellVQ, Nature Communications 17:4037 · AlphaGenome, Nature (2026) · C2S-Scale (Google Research; bioRxiv 2025.04.14.648850) · X-Cell (Xaira) · Arc Institute Virtual Cell Challenge 2025 wrap-up and 2026 announcement · Arc Virtual Cell Atlas · scBaseCount · Tahoe-100M, Cell 189(19):5945–5961 · Evo2, Nature 652:1349–1361 · Chan Zuckerberg Initiative Virtual Cells Platform, cz-benchmarks, rBio · UNAGI, Nature Biomedical Engineering (2025) · ProteinTalks, Nature (Sep 2026) · CAPTAIN, Nature Communications (2026) · VirTues, Nature (2026) · genome-wide morphology atlas, Nature Methods (2025) · FDA Modernization Act 2.0, Pub. L. 117-328 §3209 / 21 U.S.C. 355(z) · FDA NAMs draft guidance, 91 FR 13313 · FDA “Nonclinical Testing Terminology” Direct Final Rule, 91 FR 59988 · FDA AI credibility draft guidance, Docket FDA-2024-D-4689 · ICH M15, 91 FR 33179 · FDA/EMA guiding principles of good AI practice (14 Jan 2026).
Research notes: [[AI-Virtual-Cells-Cellular-Representation-2026]]
Saram Consulting