Source-linked AI summary

How to Build the Virtual Cell with Artificial Intelligence: Priorities and Opportunities

Charlotte Bunne, Yusuf Roohani, Yanay Rosen, Ankit Gupta, Xikun Zhang, Marcel Roed, Theo Alexandrov, Mohammed AlQuraishi, Patricia Brennan, Daniel B. Burkhardt, Andrea Califano, Jonah Cool, Abby F. Dernburg, Kirsty Ewing, Emily B. Fox, Matthias Haury, Amy E. Herr, Eric Horvitz, Patrick D. Hsu, Viren Jain, Gregory R. Johnson, Thomas Kalil, David R. Kelley, Shana O. Kelley, Anna Kreshuk, Tim Mitchison, Stephani Otte, Jay Shendure, Nicholas J. Sofroniew, Fabian Theis, Christina V. Theodoris, Srigokul Upadhyayula, Marc Valer, Bo Wang, Eric Xing, Serena Yeung-Levy, Marinka Zitnik, Theofanis Karaletsos, Aviv Regev, Emma Lundberg, Jure Leskovec, Stephen R. Quake

arXiv:2409.11654v2q-bio.QMcs.AIcs.LGq-bio.NC

TL;DR

Accurately modeling complex cellular systems remains difficult, especially across scales, modalities, and changing conditions. The paper proposes AI Virtual Cells that learn universal biological representations and use virtual instruments for prediction, interpretation, and in silico experimentation. Their utility depends on broad data, rigorous evaluation, and open collaborative standards, while causal discovery and data coverage remain limited.

  • Problem

    Existing approaches do not fully capture multiscale, multimodal cellular systems and their nonlinear dynamics, motivating a learned simulator of cells and tissues.

  • Method

    The paper proposes a multimodal, multiscale AIVC combining universal representations with Virtual Instruments that decode or manipulate biological states.

  • Results

    The paper presents a vision in which AIVCs simulate cellular systems, predict responses and mechanisms, and support interpretable in silico experiments across biological contexts.

  • Takeaways & Limitations

    AIVCs could help generate and prioritize testable biological hypotheses, guide experiments, and support understanding of cellular mechanisms.

  • Takeaways & Limitations

    AIVCs may not uncover phenotype causal factors through computation alone, and their evaluation must address changing distributions and broad biological generalization.

Abstract

from arXiv · show

The cell is arguably the most fundamental unit of life and is central to understanding biology. Accurate modeling of cells is important for this understanding as well as for determining the root causes of disease. Recent advances in artificial intelligence (AI), combined with the ability to generate large-scale experimental data, present novel opportunities to model cells. Here we propose a vision of leveraging advances in AI to construct virtual cells, high-fidelity simulations of cells and cellular systems under different conditions that are directly learned from biological data across measurements and scales. We discuss desired capabilities of such AI Virtual Cells, including generating universal representations of biological entities across scales, and facilitating interpretable in silico experiments to predict and understand their behavior using virtual instruments. We further address the challenges, opportunities and requirements to realize this vision including data needs, evaluation strategies, and community standards and engagement to ensure biological accuracy and broad utility. We envision a future where AI Virtual Cells help identify new drug targets, predict cellular responses to perturbations, as well as scale hypothesis exploration. With open science collaborations across the biomedical ecosystem that includes academia, philanthropy, and the biopharma and AI industries, a comprehensive predictive understanding of cell mechanisms and interactions has come into reach.

AI Virtual Cells

The paper envisions AI Virtual Cells as learned simulators that integrate biological knowledge across scales, modalities, times, and contexts. They should represent states, predict behavior and mechanisms, and support in silico experiments that guide hypotheses and data generation.

  • AI Virtual Cells: An AI Virtual Cell is a learned simulator of cells and cellular systems across differentiation, perturbation, disease, stochastic, and environmental contexts.
  • AI Virtual Cells: AI Virtual Cells should create universal representations, predict cellular dynamics and mechanisms, and perform in silico experiments that guide data collection.
  • Universal representations: Universal representations should integrate molecular, cellular, and multicellular scales while accommodating diverse modalities and contexts.
  • Predicting cell behavior and understanding mechanisms: Manipulating virtual cell states could reveal trajectories in development, homeostasis, pathogenesis, and disease progression.
  • Predicting cell behavior and understanding mechanisms: Virtual interventions could propose causal factors and reduce hypothesis space, but causal discovery may not be feasible through computation alone.
  • In silico experimentation and guiding data generation: Virtual Instruments could simulate laboratory experiments, including difficult cell systems and expensive readouts, while confidence estimates help prioritize additional data collection.

Building the AIVC

The proposed AIVC combines universal multiscale biological representations with neural Virtual Instruments that decode states or manipulate them. Its construction must balance broad molecular and cellular coverage against computational and data constraints.

  • Building the AIVC: The AIVC comprises a universal multimodal, multiscale biological state representation and Virtual Instruments that manipulate or decode it.
  • Building the AIVC: A Universal Representation is a learned embedding that preserves meaningful relationships and patterns in high-dimensional, multimodal, multiscale biological data.
  • Building universal representation across physical scales: Distinct representations can be aggregated from molecules through cells to tissues and organs, creating consistency across physical scales.
  • Molecular scale: Sequence models suit DNA, RNA, and proteins, whereas atomic-resolution models may better cover other molecules but impose greater computational and training-data demands.
  • Cellular scale: Cellular representations integrate molecular identities and quantities with locations, timestamps, imaging features, and single-cell measurements.
  • Building the AIVC: Manipulator Instruments can predict altered representations after perturbations, while Decoder Instruments produce interpretable outputs such as cell labels or synthetic images.

Data needs and requirements

Building an AIVC requires diverse, multimodal, temporally resolved data spanning biological scales and perturbations. Human-data constraints, combinatorial complexity, uncertain data requirements, and representation bias remain major obstacles.

  • Data needs and requirements: Training data should span domains, modalities, biological diversity, temporal and physical scales, and experimentally induced perturbations.
  • Data needs and requirements: Multimodal datasets linking molecular signatures, spatial organization, regulation, and cell behavior are needed to connect biological scales.
  • Data needs and requirements: Human datasets provide limited opportunities for controlled in vivo experimentation and perturbation, motivating 3D tissue systems such as organoids.
  • Data needs and requirements: Combinatorial biological spaces can exceed practical experimental and computational enumeration, requiring new exploration methods.
  • Data needs and requirements: The amount of data needed remains difficult to estimate because even one human cell system has substantial nominal complexity.
  • Data needs and requirements: Unequal representation of species, diseases, and human ancestral populations can encode biases that reduce AIVC impact.

Model evaluation

AI Virtual Cells require evaluation that tests both generalization across changing biological contexts and their ability to generate biologically meaningful discoveries. Interpretation and biological causality may also matter beyond statistical performance.

  • Model evaluation: A comprehensive, adaptable benchmark must test generalizability across biological contexts and downstream tasks under distribution shifts.Relevant shifts include environmental changes, infections, and genetic variants.
  • Model evaluation: Evaluation should prioritize both generalizability and the discovery of new biology.
  • Model evaluation: Cross-modal reconstruction can test performance in unseen cell types or genetic backgrounds by linking modalities such as morphology, gene expression, and microscopy sequences.
  • Model evaluation: Biological relevance can be assessed through testable hypotheses and experimentally verifiable phenotypes such as growth rates or molecular profiles.
  • Model evaluation: Interpretability and biological causality may be required alongside statistical performance as AI Virtual Cell capabilities improve.

Interpretability and interaction

The paper argues that AI Virtual Cells should become interpretable, interactive systems rather than opaque predictors. Mechanistic explanations and accessible interfaces could help researchers investigate predictions and use them across expertise levels.

  • Interpretability and interaction: Fully mechanistic models may be sacrificed for data-learned interactions that generalize beyond observations, making increased interpretability especially desirable.
  • Interpretability and interaction: Multi-scale interactions and modular structure could connect predictions to disrupted genes, proteins, or molecular processes involved in disease.
  • Interpretability and interaction: An interactive layer should help researchers with varying expertise grasp and use AI Virtual Cell predictions.
  • Interpretability and interaction: Large-language-model agents could serve as virtual research assistants and use scientific literature to provide deeper insight into AI Virtual Cell predictions.

An open collaborative approach

Building AI Virtual Cells is presented as a large, long-term undertaking requiring open infrastructure, shared resources, and collaboration across scientific and industrial communities. Accessibility, diversity, and privacy are central design considerations.

  • An open collaborative approach: Developing an AI Virtual Cell requires substantial investment, diverse expertise, many iterations, and a concerted open-science effort.
  • An open collaborative approach: Open data resources, data standards, collaborative modeling platforms, and benchmark datasets can support accessible and coordinated development.
  • An open collaborative approach: Infrastructure should represent human ancestral and geographic diversity while safeguarding privacy and supporting affordable model access.
  • An open collaborative approach: A shared platform should connect biologists, clinicians, and computer scientists, enable lab-to-model iteration, and support rapid testing and benchmarking.
  • An open collaborative approach: Successful interactive AI Virtual Cell models could change how cell biology research is conducted and serve as open hubs for collaboration, deployment, and education.

Outlook and reasons for optimism

The paper envisions AI Virtual Cells as an open interface between large biological datasets, computational experimentation, and engineered biology. Their potential extends from unified biological understanding to drug discovery, synthetic biology, and broader scientific exploration.

  • Outlook and reasons for optimism: Large reference projects have made extensive genomic, transcriptomic, proteomic, cellular, and population-scale data available for training machine-learning models.
  • Outlook and reasons for optimism: AI Virtual Cells could function as virtual laboratories linking in silico experiments with physical-laboratory results and supporting biomedical research, personalized medicine, drug discovery, and cell engineering.
  • Outlook and reasons for optimism: By connecting generative AI, AI agents, and biology, AI Virtual Cells could help scientists understand cells as information-processing systems and design novel synthetic biology.
  • Outlook and reasons for optimism: Open sharing of data, models, benchmarks, and contextualized findings is proposed to foster continual improvement across the scientific community.
  • Outlook and reasons for optimism: The convergence of AI and biology is presented as a potential paradigm shift for exploring cellular mysteries through safe, ethical, and reliable AI.

Competing interests

The paper outlines challenges and opportunities for AI Virtual Cells, emphasizing interpretability, collaboration, ethical data use, and applications in disease modeling and therapy.

  • Challenges and requirements: AI Virtual Cells must balance highly accurate, calibrated biological predictions with interpretability and actionable outputs for experimental validation.Proposed explanatory approaches include causal modeling, sparse featurization, counterfactual reasoning, and interfaces supported by AI research agents.
  • Challenges and requirements: Developing AI Virtual Cells requires open collaboration, shared infrastructure, and engagement across researchers, educators, patients, and the public.The paper envisions interconnected platforms supporting collaborative model development, deployment, and training.
  • Challenges and requirements: Open datasets representing human diversity must be developed with ethical transparency and safeguards against falsified or contaminated data.The authors identify responsible data use and contamination mitigation as substantial challenges.
  • Applications: Disease-focused AI Virtual Cells could integrate patient and disease-context data to test interventions, prioritize virtual hits, and support phenotypic screening.The paper states that prioritizing virtual hits with higher chances of success could lower experimentation costs and accelerate development.
  • Applications: AI Virtual Cells could support individualized cell engineering and precision oncology by modeling patient variation, cellular interactions, and tumor microenvironment structure.Examples include engineering pancreatic beta cells, identifying pan-cancer tumor niches, and modeling patient-specific cancer biology.

A hypothesis-generating framework for scientific research

The paper proposes AI Virtual Cells as hypothesis-generating systems that computationally explore many possible biological interventions and identify informative experiments. This approach shifts computational modeling beyond analyzing prior experiments toward iterative in silico experimentation and model improvement.

  • A hypothesis-generating framework for scientific research: Virtual Cells could computationally explore vast arrays of biological hypotheses rather than only analyze data through an existing hypothesis.The proposed paradigm uses in silico experimentation to expand the range of hypotheses considered.
  • A hypothesis-generating framework for scientific research: Virtual Cells could identify the most informative experiments for addressing specific biological questions.This reframes computational models as tools for selecting experiments, not merely validating hypotheses or processing observations.
  • AI techniques for building the AI Virtual Cell: AI Virtual Cell development requires connecting diverse neural architectures while weighing trade-offs in accuracy, speed, and generalizability.The paper presents architecture selection as dependent on biological modality and inductive bias rather than a single universal model type.

Transformers

Transformers process biological tokens by integrating contextual information through self-attention, with positional encodings added when sequence order matters. In cellular applications, tokens can represent molecules or genes and model their interactions.

  • Transformers: Transformer layers process token sequences through self-attention followed by feed-forward networks to produce context-enriched representations.Tokens may represent words, RNA molecules, or gene representations.
  • Transformers: Self-attention can model gene interactions when RNA molecules detected by single-cell sequencing are represented as tokens.This applies the transformer’s context-integration mechanism as a biological inductive bias.
  • Transformers: Positional encodings enable transformers to represent sequence-specific dependencies in DNA and other biological sequences.They support applications such as masked language modeling, where missing sequence tokens are predicted from context.

Convolutional neural networks

Convolutional neural networks learn hierarchical spatial features through convolution, pooling, and downstream interpretation layers. In biology, they support image-based analysis of cells, tissues, and molecular patterns, while sequence applications and vision transformers extend their use.

  • Convolutional neural networks: Convolutional neural networks learn spatial feature hierarchies through convolutional filters, pooling, and fully connected layers.The architecture is primarily used for analyzing images.
  • Convolutional neural networks: CNNs support multiplex imaging by detecting complex patterns and structures across labeled molecular targets and heterogeneous tissues.This makes them useful for studying interactions among molecules or cell types in tissue environments.
  • Convolutional neural networks: CNNs can model DNA sequences, while vision transformers increasingly supplement or replace them for image tasks requiring global context.The paper contrasts local convolutional processing with self-attention over entire images.

Diffusion models

Diffusion models generate structured biological data by transforming random noise, while flow matching extends this framework to continuous cellular transformations over time and space.

  • Diffusion models transform random noise into structured outputs such as images, text, and cellular states.They are generative deep learning models designed to produce high-quality, diverse samples.
  • Their distribution-learning and temporal-spatial modeling capabilities suit the high-dimensional, intricate data structures found in biological systems.

Graph neural networks

Graph neural networks model biological systems as graphs, using relationships between connected nodes to represent structures such as proteins and spatially organized tissues.

  • Graph neural networks model graphical data using nodes connected by edges.In biological applications, graphs can represent protein structures or spatial relationships among cells.
  • Protein residues can form nodes linked by bonds, while neighboring cells in a tissue can form nodes linked by physical proximity.For spatially organized cells, these connections can represent potential chemical signaling relationships.
  • At each layer, a node updates its representation using its current representation and those of neighboring nodes.Stacking layers lets nodes receive information from increasingly distant neighbors through multiple hops.
Loading 2409.11654v2…