Get all your news in one place.
100's of premium titles.
One app.
Start reading
Medical Daily
Medical Daily
Ryan Archer

Biology Has No Flight Simulator, and Three Researchers Just Published the Blueprint for One

Drug development runs on a brutal arithmetic. Most candidates fail, and they usually fail late, after the money is spent. A proposal published this week argues that the way out is to stop treating biology as something you can only poke at in a dish and start treating it as something you can simulate.

Le Song, Eran Segal, and Eric Xing laid out the engineering blueprint in a Nature Medicine Perspective, How to build an AI-driven digital organism.

The comparison the field reaches for is aviation. Pilots do not learn crosswind landings by crashing airliners. They learn in simulators. Manipulating biology, the authors argue, is complex, expensive, and risky enough that it should be preceded by extensive computer-aided design and simulation, the way civil, nuclear, and semiconductor engineering already are.

The Argument Is That One Giant Model Will Not Work

Most AI systems in biology today do one thing well. AlphaFold predicts protein structures. Other models classify cell types or predict how gene expression shifts after a perturbation. Each is narrow by design.

The perspective argues that the leap to a digital organism is not a bigger version of any of those. It is an architecture problem. The proposed system is a modular collection of foundation models spanning DNA, RNA, proteins, structures, and single cells, connected to reflect biological scales rather than stacked arbitrarily, and the paper sets out a three-stage roadmap for assembling one.

The claimed payoff is a platform that would allow a researcher to alter a gene or introduce a compound and see the predicted downstream consequences before committing to a bench experiment, thereby narrowing the field of hypotheses worth testing. The reason the architecture matters is that biology is not a single problem at a single scale. A change in a DNA sequence propagates upward through RNA, protein, structure, network and phenotype, and a model that handles only one of those layers cannot follow the consequence.

One disclosure belongs alongside that claim. All three authors declare a financial interest in GenBio AI, the company building the system they describe. The paper was peer-reviewed, but it is a vision statement from people with a stake in the vision.

This Is a Crowded Race with No Winner Yet

The idea is not one company's. A Cell paper from a large, multi-institutional group published in 2024 set out priorities and design principles for an AI virtual cell, and the concept has since become one of the most heavily resourced ambitions in computational biology, with the Chan Zuckerberg Initiative and the Arc Institute among the groups working toward it.

A more contained version appeared in Nature as a perspective on a virtual yeast. That proposal decomposes cellular complexity into eight function-centered modules spanning genetic, metabolic, and structural systems, each realized as a domain-specific AI tool and coordinated through a large language model orchestration layer, with a closed-loop pipeline that designs and runs its own experiments. Baker's yeast is genetically tractable and data-rich, making it a plausible proving ground in a way that human cells are not yet. It also has decades of systematic gene-deletion and gene-interaction data, which is exactly the sort of training material a simulator needs and that human tissue largely lacks.

The Obstacles Are Not Small, and the Field Knows It

A Nature news feature on virtual cells published earlier this year captured the central difficulty plainly: researchers are still working out how to reproduce life's complexity without drowning in data.

Integrating fundamentally different data types remains an unsolved problem. Model outputs are frequently uninterpretable. Computational demands are steep. And there is a subtler danger in the training data itself: batch effects can cause the fingerprint of a particular lab, reagent kit, and day to leak into a model and be learned as if it were biology. A confident prediction built on a batch artifact is worse than no prediction, because it looks trustworthy.

There is also a harder test that the field has largely failed to pass. A useful virtual cell must predict responses to perturbations it has never seen, not interpolate between measurements it was trained on. And the same drug behaves differently across cell contexts, which any general-purpose simulator must reproduce rather than average away.

What This Is Not

Nothing described here is a virtual human. It is not a replacement for clinical trials, and it does not yet reliably predict how a drug will affect a patient.

What is being proposed is an engineering roadmap for a research tool, published as a perspective rather than as validated results. Its authors are candid that the components exist in pieces and the integration does not. The earlier version of this argument circulated as an arXiv preprint two years ago, and what has changed since is the peer review, not a working simulator.

That distinction matters for anyone reading about it. The plausible near-term benefit is upstream: fewer wasted experiments, better-prioritized candidates, faster elimination of ideas that were never going to work. Predictions from any such system will still have to be confirmed in living cells and, eventually, in people before they change anything a patient receives.

Key Questions Answered

What is an AI digital organism?

A computational system intended to simulate biology across scales, from molecules to cells to individuals, so researchers can predict the effect of a drug or genetic change before running the physical experiment.

What did the new paper actually publish?

An engineering blueprint, not a working simulator. It describes a three-stage roadmap for building component AI models and connecting them across biological scales.

How is this different from AlphaFold?

AlphaFold solves one narrowly defined problem: protein structure. The proposed system aims to link many specialized models, spanning from DNA to whole organisms, into a single platform.

Do the authors have a conflict of interest?

Yes, and it is disclosed. All three declare a financial interest in GenBio AI, the company developing the system described in the paper.

What are the main technical risks?

Data integration across modalities, uninterpretable outputs, heavy computing demands, and batch effects, in which laboratory artifacts get learned as if they were real biology.

Could this replace clinical trials?

No. It is a research prioritization tool. Any prediction would still require validation in living systems and in patients before it could affect medical care.

Sign up to read this article
Read news from 100's of titles, curated specifically for you.
Already a member? Sign in here
Related Stories
Top stories on inkl right now
One subscription that gives you access to news from hundreds of sites
Already a member? Sign in here
Our Picks
Fourteen days free
Download the app
One app. One membership.
100+ trusted global sources.