Virtual cells: Perspectives from structure prediction
What experiments and concepts map onto cellular function?
The virtual cell, or a computational model of the cell, is a great idea on paper. Unfortunately, the catchiness of the term “virtual cell” has been more successful than any result produced to date. Today, the term lacks any real specificity, despite the hype and the amount of resources spent to make it a reality. Underlying this problem is the fact that no one seems to agree on what abstraction best represents a cell’s function and what experimental data best captures it.
The enthusiasm is easy to grasp. It stems from a mix of two forces: genuine wonder for biology and the belief that AI models can solve complicated problems if presented with enough data. For example, Jensen Huang, CEO and chairman of NVIDIA, voiced this curiosity, wonder, and enthusiasm in a recent appearance when he asked, “Why does a cell do what a cell does?” Later, he mentioned that we have to figure out how to best represent this data to computers to extract valuable insights. In aggregate, these statements sum up the crux of the virtual cell project.
Putting implementation aside for the moment, the ideal virtual cell is a computational simulation that predicts what a cell will do when initialized with a set of conditions. Suspend your disbelief and imagine that we could set up a cell with a specified configuration of molecules. If that system begins out of equilibrium, we could watch it relax — from a perturbed state to steady state. A simulation that captured these dynamics would let us interrogate a wide range of cellular properties, including the cell’s intrinsic behaviors during disease and therapeutic interventions. The following questions all fall naturally within its purview and resemble experiments performed today in research labs:
What will a cell do if you add, remove, or mutate gene X?
What happens if the genome is shuffled? What are the alternate wirings?
Where are the equilibrium positions of the various organelles?
What are the limits of the metabolic states a cell can tolerate?
What cell types and cell behaviors does a cell’s genome encode?
What physical shapes can a cell take?
And there are many more…
Despite the seemingly endless possibilities, there is growing frustration with the virtual cell as a brand. That frustration stems from a mismatch between the ambition — to simulate and design life — and our current ability to harness and integrate observations from fields as disparate as cell biology, metabolism, gene regulation, biochemistry, and signaling. Spend time with anyone who thinks hard about how post-translational modifications (the addition of small sets of atoms to a protein) dictate cellular behavior or how cell state dictates mitochondrial behavior, and they will tell you we do not know enough to model those individual properties in any cell, let alone jointly alongside every other cellular process. To make matters worse, disparate data formats make it difficult to synthesize the data we do have into a coherent framework.
Lessons from protein structure prediction
Why, then, do I think constructing a virtual cell is possible despite its grandiosity? My intuition comes primarily from AlphaFold, which I consider the most successful molecular simulator we have. In today’s parlance, efforts to generate AlphaFold-like models may be called the “Virtual Protein Project”. These models predict the structures of unseen proteins and complexes, and provide insight into the effects of mutations on structure. All of this is done from protein sequences directly. Give these models any sequence, and they will output a predicted structure and a calibrated confidence metric. These two outputs allow computational biologists to perform in silico experiments on proteins that have never been synthesized in a lab and perform dry runs of wet-lab experiments to rule out those that are unlikely to work, saving precious time and resources. Although there is plenty of room for improvement, including training models that understand atomic dynamics, explicit solvent, energetics, and molecular orbitals, the conceptual framework for training, evaluating, and deploying such models is now well established.
This conceptual framework has also allowed us to assess protein designs. The protein design problem is the inverse of the structure prediction problem. Given a desired 3-dimensional protein shape, what amino acid sequence or sequences will fold into it? A naive and inefficient method to design a 3-dimensional protein structure would be to guess a sequence at random, run it through AlphaFold and check the similarity between the desired structure and the predicted structure. Using sequences generated from more sophisticated design methods, models like AlphaFold enable protein design by evaluating novel designed sequences. A huge benefit to having a clear conceptual framework is that it facilitates defining and solving this inverse problem.
A similar framework could allow us to build a useful virtual cell – though it’s worth stating explicitly that we have not yet outlined an analogously well-defined task and metric for such a model. What’s the cellular equivalent of a protein’s amino acid sequence? What exactly does the virtual cell predict? And at what spatial and temporal resolution? Defining this task is of paramount importance for a virtual cell project. And assessing the relationship between structure and function will be just as important.
Structure → Function
Structural biology has another nice property: when you look at a protein’s structure, the positioning of its atoms in 3-dimensional space, you can get a sense of what that protein is capable of. The positions of the atoms on a protein’s surface show you how it might bind to other proteins. How the atoms are arranged in an active site tells you how it organizes atoms to catalyze a chemical reaction. Because structure and function correspond, every structure that is solved confirms some previously discovered aspect of that protein’s biochemistry or biophysics. And when a structure is solved for an unknown protein, one can form hypotheses based on its organization alone: the structure could reveal the surfaces of a viral fusion protein exposed to immune recognition or facilitate the search for other, similar structures that are well-characterized (e.g., similar enzyme active sites). Structures can also help explain why certain mutations cause disease, while others don’t.
What is the analogous structure/function measurement for the cell? Intuitively, watching a cell tells you a great deal about cellular function. When you see one cell eating another, that process is easy to describe in natural language. Similarly, when a cell is decorated with spindle-like arms, branching like the roots of a tree, one can infer that its function is to make connections, or to spread out and sense. As with protein folds, many different cell types can converge on the same properties with very different underlying genome sequences and gene-network topologies. Before the advent of molecular measurements, these visual properties and the functional descriptions that accompanied them dominated the life sciences.
Neurons were famously described by Ramón y Cajal, who illustrated them beautifully in the late 1800s and early 1900s. Working only with Golgi’s silver stain, Cajal argued two things from the neuron’s morphology alone: 1) that neurons are discrete cells rather than a continuous brain reticulum, and 2) that signal flows through neurons directionally. At the time, Cajal had no access to the molecules involved. The synapse would not be named for another decade, and chemical neurotransmission would not be demonstrated for thirty years. Looking at the shape of the cell was enough to describe and intuitively understand putative functions.

However, this is often easier said than done. Cajal used a stain which turned roughly 1% of the cells in a tissue black, revealing the striking shape of individual neurons. Today, tools like electron microscopes and fluorescent tags for proteins allow us to resolve much more, including the molecular composition of cells. But much like solving an individual protein’s structure, imaging cells at this high resolution in a tissue is difficult to establish and even more difficult to generalize. Worse, unlike proteins, which have a single data format for describing their structure (PDB or mmCIF), every imaging experiment has its own format, using different optics, detectors, and contrast agents. The result is a heterogeneous, distributed knowledge base, rather than one that is standardized and centralized.
Finally, it is unclear whether having a snapshot of the 3-dimensional positions of proteins within a cell would inform cellular function the way a protein structure informs protein function. Unlike the strong correspondence of the atom’s position in a protein and its corresponding function, proteins in a cell function through their dynamic movement within cells. Proteins tend to occupy territories within the cell, but also explore different regions to perform their function. Perhaps the functional analog to a protein structure would be a cellular video that shows the movements of many proteins within the cell simultaneously. Together these technical and conceptual considerations mean that data from imaging studies will likely require more technology development, and thus proceed at a slower pace.

Sequencing proxies for function
In the absence of new imaging tools, a number of public, private, and philanthropic efforts are attempting to generate data to train a virtual cell model. Organizations such as Xaira, Biohub, and Tahoe Therapeutics all have programs to perturb cells genetically, chemically, or metabolically, and read the outcome using single-cell RNA sequencing. Unlike imaging cells in a tissue, this is a procedure that can be performed on real samples – cells taken directly from a person or an organism. It generates a standardized output, the count matrix, and reveals aspects of cellular heterogeneity that correspond to biological functions of interest: cell type composition, cell state shifts, and gene expression changes. And unlike imaging, it scales.
Unfortunately, the count matrix is woefully incomplete as a measure of cellular state. It does not even contain every transcript a cell expresses, excluding microRNAs, transfer RNAs, and anything else without a polyA tail. More importantly, aspects of a cell that matter for function, such as protein phosphorylation, oxidative state, morphology, and tissue position, are poorly represented in the transcriptome. Without a coherent organizing principle for how a cell transcriptome captures cell function, these data collection efforts amount to applying sensitive tools to a problem they are not meant to solve.
If we look to the analog in structural biology, the quest to predict structure from sequence alone has been fruitful. Many lines of evidence indicate that the main driver of AlphaFold’s success is the identification of co-evolving residues, made apparent in sequence alignments, that act as constraints on protein structure. Statistical methods such as direct coupling analysis and deep learning models such as ESM exploit this signal by reading alignments directly or by absorbing the same statistics from pretraining across billions of sequences.
A path forward
So how do we move towards a virtual cell?
Following the analogy of protein structure prediction, virtual cell models must draw on comparable evolutionary signals to constrain predictions of cellular function. These constraints must incorporate domain-knowledge to specify appropriate model topologies (like AlphaFold’s ingenious triangle-update), while simultaneously drawing upon the rich evolutionary data spanning the tree of life. For example, the gains and losses of cellular traits, like heritable features running through a family tree, map to changes in genes and could therefore act as approximate constraints for a cell or organism. The genes to construct a nuclear membrane, nuclear pore and nuclear cytoskeleton were likely preconditions for the formation of the nucleus during eukaryagenesis. Similarly, matching the flux of genes in organisms across the tree of life with the flux of cellular traits should facilitate building a map from genome sequence to cellular function.
However, while sequencing genomes is easy, observing and pinpointing the gain and loss of traits is hard. Data linking genotype to phenotype come primarily from two sources: nationalized genome programs that pair medical records with genome sequences, and decades-long trait mapping efforts in model organisms. Every human genome may well be sequenced in the near future, but a billion samples from one species still only spans a narrow slice of biology – far too narrow to delineate different cell functions. Domesticating new organisms and onboarding genetic and imaging tools for each of them would widen that slice, but every organism demands the development of bespoke methods and years of effort.
Despite these headwinds, developing new tools for gathering cellular data, training AI models, and interpreting their outputs must follow a coherent set of organizing principles. Otherwise cellular modeling will go the way of all-atom molecular dynamics, adrift in a sea of molecular details, unable to say which ones matter and solving problems that may not yield general biological principles. Our perspective is that cellular evolution has already run a massive experiment, constraining the components, topologies and systems required for cellular function. To that end, our organizing principles are as follows: to read the results of this evolutionary process at scale, match the flux of genes to the flux of traits across the tree of life, and build the computational tools that make that variation legible in terms of cellular function. This will (hopefully) result in a cellular equivalent to AlphaFold – a virtual model that can reliably predict and eventually design biology in its native language.




