Source-linked AI summary
$D^{2}R^{2}$: Discrete Diffusion with Regulation Reinforcement for Single-Cell Perturbation Prediction
Ninghan Fan, Qi Liu, Xunuo Zhu, Yukai Sun, Luyuan Chen, Xuheng Zhou, Yuetian Du, Ming Kong, Xiaojun Zhu, Jie Liu, Zhan Zhou, Qiang Zhu
TL;DR
Single-cell perturbation prediction remains difficult because of complex regulatory mechanisms, biological heterogeneity, and dynamic cellular responses. D2R2 addresses this by progressively generating gene-wise expression tokens under an adaptive, biologically grounded ordering policy. It ranks first across all metrics on Norman19 and leads in most metrics on H1, while biologically informed ordering outperforms random and uncertainty-based strategies.
Problem
Single-cell perturbation prediction remains difficult because gene regulatory mechanisms, biological heterogeneity, and cellular responses are complex and dynamic.
Method
D2R2 progressively generates masked expression tokens while a regulatory policy initializes gene ordering from control-derived networks and adapts it through reinforcement learning.
Results
D2R2 ranks first across all metrics on Norman19 and leads in most metrics on H1, while biological-prior ordering outperforms random and uncertainty-based strategies.
Takeaways & Limitations
Gene generation order is an effective and biologically interpretable component of single-cell perturbation prediction.
Abstract
from arXiv · showhide
Predicting single-cell transcriptomic responses to genetic perturbations is central to functional genomics and virtual-cell modeling. Existing approaches, however, typically predict an entire expression profile as a whole, leaving the order in which individual gene responses are generated unmodeled. To address this problem, we introduce \textbf{$D^{2}R^{2}$} (\textbf{D}iscrete \textbf{D}iffusion with \textbf{R}egulation \textbf{R}einforcement), which reformulates perturbation prediction as regulation-guided gene-wise progressive generation. A Masked Discrete Diffusion Model represents expression as ordinal tokens and reconstructs a fully masked profile step by step, allowing generated gene responses to condition those that remain masked. A Regulatory Policy Module initializes the generation policy from a gene regulatory network inferred from control cells and adapts it to the perturbation and current partially generated state. Then, group-relative policy optimization refines only the ordering policy using final perturbation-effect agreement as reward. Across Norman19 and VCC-H1, $D^{2}R^{2}$ achieves the best performance on all five metrics on Norman19 and remains competitive on H1. Controlled ablations holding the generator and generation budget fixed show that biological-prior ordering improves over random ordering and is more reliable than uncertainty-based heuristics, whereas reversing the biological-prior ordering degrades every metric. Biological analyses further show that the refined policy prioritizes regulatory genes early while promoting perturbation-specific transcription factors and responsive genes. These results establish gene generation order as an effective, controllable, and biologically interpretable dimension of single-cell perturbation prediction.
Introduction
D2R2 reformulates single-cell perturbation prediction as gene-wise progressive generation, using regulation-guided adaptive ordering with discrete diffusion. Experiments indicate that biological-prior ordering improves prediction and that refined orders preserve regulatory prioritization while adapting to perturbations and generation state.
- Motivation: Perturbation-response prediction remains difficult because of complex gene regulatory mechanisms, biological heterogeneity, dynamic cellular responses, and costly experimental data generation.These challenges motivate virtual-cell models for simulating cellular responses in silico.
- Problem formulation: The approach makes gene-generation order explicit, allowing resolved responses to condition genes that remain masked during progressive expression generation.This reformulation addresses the limitation of profile-level methods that generate the full expression profile at once.
- Method: D2R2 couples a Masked Discrete Diffusion Model with a Regulatory Policy Module that initializes regulatory-network ordering and refines it using GRPO and perturbation-response agreement.The generator remains fixed while the policy becomes state- and perturbation-specific.
- Empirical validation: Across Norman19 and VCC-H1, biological-prior ordering outperforms random and uncertainty-based strategies, while reversing the prior consistently degrades prediction.Controlled ablations hold the MDDM fixed, showing that generation order is consequential.
- Biological analysis: The refined policy preserves early regulatory-gene prioritization while promoting perturbation-specific transcription factors and responsive genes.This provides a biologically interpretable characterization of the learned generation orders.
Related Work
Prior single-cell perturbation predictors span latent-variable, disentanglement, distributional, and flow-based formulations, while other work incorporates regulatory structure into gene representations and prediction. Reinforcement learning has also been applied to biological design problems where local decisions are evaluated through delayed objectives.
- Predictive and generative formulations: Single-cell perturbation prediction includes latent-space shift models (Lotfollahi, Wolf, and Theis 2019), disentanglement-based predictors, and distributional or flow-based approaches.These formulations differ in how they represent perturbation effects and mappings between control and perturbed cell populations.
- Regulatory structure: Regulatory information from gene networks, transcription factor-target relationships, and pathways has been incorporated through graph neural networks, message passing, structure-aware embeddings, and adjacency regularization (Adduri et al. 2025; Theodoris et al. 2023).Graph-based perturbation models also use gene–gene relationships to improve response prediction (Roohani, Huang, and Leskovec 2024), but these approaches primarily shape representations or inputs.
- Reinforcement learning: Reinforcement learning supports computational-biology problems with difficult-to-supervise local decisions but evaluable complete designs, including molecular optimization, molecular-graph construction (You et al. 2018), RNA inverse folding (Runge et al. 2018), and protein design.This line of work motivates sequential-generation settings with delayed objectives.
Method
D2R2 separates expression-token generation from gene-generation ordering: a masked discrete diffusion model predicts unresolved responses, while a regulatory policy selects which genes to resolve next. The policy is initialized from a control-cell regulatory network and refined with perturbation-effect rewards while the generator remains fixed.
- Generation order: The generation-order strategy assigns priorities within the current masked set and selects a subset of genes to resolve at each iteration, while keeping predicted tokens fixed thereafter.This separates expression modeling from ordering and permits comparisons among random, confidence-based, and entropy-based strategies with the generator held fixed.
- Masked discrete diffusion: D2R2 tokenizes continuous expression into ordinal bins and uses a masked discrete diffusion model to predict original tokens at unresolved gene positions, conditioning on control, perturbation, and cellular context.A special mask token explicitly identifies unresolved genes, allowing generated genes to provide context for masked genes; the model follows masked diffusion formulations (Austin et al. 2021; Nie et al. 2025a,b).
- Biological-prior ordering: D2R2 initializes ordering from a cell-type-specific regulatory network inferred with DeepSEM (Shu et al. 2021), ranking genes by PageRank so regulatory genes are generated early and contextualize later predictions.The biological prior is restricted to unresolved genes at each step before initializing the policy.
- Regulatory Policy Module: The Regulatory Policy Module converts the static prior into a perturbation- and state-specific policy by encoding the partially generated state, perturbation, cellular context, and unresolved genes.RPM recomputes priorities after each update, and inference selects the next subset deterministically with Top-K under a predefined generation schedule.
- Policy refinement: After biological-prior initialization, the MDDM and prior are frozen while RPM is refined with GRPO (Shao et al. 2024) using group-normalized perturbation-effect correlation rewards.Training samples gene subsets without replacement, resolves them with the MDDM, completes trajectories with a frozen reference policy, and uses deterministic Top-K selection at inference.
Experiments
D2R2 ranks first across all five metrics on Norman19 and leads most metrics on H1, including all three differential-expression metrics on both benchmarks. Controlled ablations and biological analyses show that regulatory ordering and reinforcement learning produce effective, perturbation-specific generation priorities.
- Experimental setup: The evaluation uses Norman19’s held-out-combination split and H1’s held-out-cell split, with matched preprocessing, splits, and evaluation across comparison methods.Norman19 contains 287 perturbation conditions, including 131 combinatorial perturbations, while H1 contains approximately 400,000 cells across 300 CRISPRi targets.
- Main results: D2R2 ranks first across all metrics on Norman19 and leads most metrics on H1, including all three differential-expression metrics on both benchmarks.It achieves the best centroid agreement for absolute expression on Norman19 and remains competitive on H1.
- Generation-order ablation: Uncertainty-based ordering provides little guidance because gains on some metrics often coincide with substantial deterioration on others.Prioritizing low-confidence genes improves Pearson ∆ and Cos. PCA but worsens Cos. LogFC and perturbation discriminability.
- Generation-order ablation: Biological-prior ordering improves most metrics over Random, whereas reversing the prior order causes every metric to fall below Random.All configurations use the same frozen MDDM and generation budget, isolating the effect of gene selection order.
- Biological analysis: RPM enriches early generation steps for regulatory genes, prioritizes perturbation-specific transcription factors, and promotes genes capturing most perturbation-specific DE genes.With biological initialization, RPM improves all three perturbation-effect metrics, while RPM-promoted genes are also enriched in cell-cycle-related pathways.
Conclusion
D2R2 makes gene generation order explicit for single-cell perturbation prediction by combining progressive masked-token generation with biologically initialized, perturbation-adaptive ordering. It ranks first across all Norman19 metrics and leads in most H1 metrics, while ablations and biological analyses support its ordering strategy.
- Conclusion: D2R2 combines progressive masked expression-token generation with a reinforcement-learned ordering policy initialized from biological priors and adapted to each perturbation.MDDM progressively generates masked tokens, while RPM adapts the ordering policy through reinforcement learning.
- Conclusion: D2R2 ranks first across all metrics on Norman19 and leads in most metrics on H1.
- Conclusion: Controlled ablations show that biologically informed ordering outperforms random and uncertainty-based strategies.
- Conclusion: RPM preserves upstream regulatory context while prioritizing perturbation-specific regulators and responsive genes.
Appendix Background of Virtual Cell
Virtual cell models aim to represent and predict cellular behavior across molecular, cellular, and tissue scales using multimodal biological measurements. Despite expanding perturbation datasets and AI methods, reliable virtual cells remain challenging because biology is nonlinear, heterogeneous, context-dependent, and mechanistic interpretability is limited.
- Motivation: Virtual cell modeling seeks predictive and mechanistic representations that simulate, predict, and steer cellular responses to perturbations across biological scales.The broader vision includes modeling transitions across diverse cellular contexts and emergent phenotypes.
- Virtual Cell Models: Artificial Intelligence Virtual Cells integrate multimodal measurements, including genomic sequences, transcriptomics, proteomics, imaging, and spatial omics, into representations spanning molecular, cellular, and tissue scales (Bunne et al. 2024).The framework defines AIVCs as multi-scale foundation models for learning universal biological representations.
- Experimental Foundations: High-throughput technologies, including CRISPR-based perturb-seq, perturbation proteomics, and single-cell imaging, provide paired pre- and post-perturbation observations for learning cellular response dynamics (Dixit et al. 2016).These experiments measure cellular states at unprecedented scale and resolution.
- Challenges: AI-driven approaches commonly predict post-perturbation transcriptomic states from initial cell states and perturbation conditions, but reliable virtual cells remain limited by nonlinear regulation, network rewiring, heterogeneity, context dependence, and weak mechanistic grounding.scRNA-seq is central because gene-expression profiles provide scalable, information-rich representations of cellular states.
Experimental Details · Datasets, Preprocessing, and Splits
Experiments use fixed 1,000-gene representations for Norman19 and H1, with training-fitted binning and control-cell references defining preprocessing. Evaluation uses fixed held-out splits: combination prediction for Norman19 and the benchmark’s saved held-out-cell split for H1.
- Datasets, Preprocessing, and Splits: Both datasets retain the top 1,000 highly variable genes, whose identities remain fixed during tokenization and generation.
- Datasets, Preprocessing, and Splits: Binning statistics are fitted on training data and reused for validation and test examples, while control cells provide reference expression for perturbation-induced changes.
- Datasets, Preprocessing, and Splits: Norman19 includes 287 perturbation conditions, including 131 combinatorial conditions, and uses a seed-42 combination-prediction split with held-out combinations reserved for evaluation.
- Datasets, Preprocessing, and Splits: H1 uses the benchmark’s saved held-out-cell split; splits remain fixed across repeated runs, and test data are used only for final evaluation.
Evaluation Metrics
Evaluation follows the PerturBench protocol, computing metrics per perturbation condition and averaging them across evaluation conditions. The metrics compare predicted and observed perturbation effects relative to control using correlation, cosine similarity, ranking, and PCA-based representations.
- Evaluation Metrics: Metrics are computed for each perturbation condition and then averaged across evaluation conditions, using predicted, observed, and control mean-expression profiles.The profiles are denoted ˆy_p, y_p, and y_ctrl, respectively.
- Evaluation Metrics: Pearson Δ correlates predicted and observed perturbation effects after subtracting the control profile, while Cos. LogFC compares relative log2 fold-change vectors using a 0.1 pseudocount.Cos. LogFC Rank evaluates perturbation-level discrimination by ranking the matched observed perturbation against other conditions; lower values are better.
- Evaluation Metrics: Cos. PCA measures cosine similarity between predicted and observed condition centroids after projection in PCA space.
Baseline Protocol
The comparison protocol standardizes data, splits, prediction space, output schema, controls, evaluation space, aggregation, and metrics across methods while preserving each model’s released representations and training procedures.
- Comparison protocol: All methods use the same 1000-gene prediction space, annotations, data partitions, held-out test conditions, AnnData schema, reference controls, PCA space, aggregation, and metrics.This design isolates differences in model design rather than data or evaluator.
- Implementation protocol: Methods retain model-specific input representations and training procedures, using unified interfaces for integrated methods and only input/output adaptations for standalone methods.Integrated methods include SAMS-VAE, Biolord, and Cell Flow; standalone methods include STATE, Squidiff, and scDFM.
D2R2 Configuration … The Results of GSEA Analysis
D2R2 uses a masked discrete diffusion generator with fixed inference and optimization protocols, while regulatory analyses examine generation order through GRN categories, TF specificity, differential expression, and pathway enrichment. GSEA shows that RPM-promoted genes are enriched in cell-cycle-related pathways, including cell-cycle and G2-M transition pathways.
- D2R2 Configuration: D2R2 uses 50 expression bins and a 12-layer Transformer with hidden dimension 768, 12 attention heads, SwiGLU dimension 3072, RMSNorm, and 0.1 dropout.Norman19 uses 64 control tokens and two perturbation tokens, whereas H1 uses one control token and four perturbation tokens.
- D2R2 Configuration: At inference, generation runs for 20 steps, updating 50 genes per step with an MDDM sampling temperature of 1.0.The denoising predictor is trained with a masked-token objective, and the dataset-specific learning rate is selected on the validation split.
- D2R2 Configuration: After MDDM training, its parameters remain fixed while RPM is optimized with GRPO using four rollouts per example and one clipped policy update from group-relative advantages.The bin-representative-value analysis found that posterior expectation had similar or slightly better mean-profile performance but substantially degraded MMD-PCA and Sym. KL.
- Repeated Runs and Statistical Protocol: Results use fixed splits, five independent seeds, and mean ± sample standard deviation over runs, with D2R2’s RPM independently initialized while its MDDM checkpoint remains fixed.Lower-is-better metrics are sign-reversed for paired improvements, and these condition-level intervals differ from training-seed standard deviations.
- Compute Environment: Experiments used Ubuntu 22.04, Python 3.10, CUDA 12.4, PyTorch 2.6, Lightning 2.6, Scanpy 1.11, NVIDIA A100 GPUs, and bfloat16 mixed precision.Metric computation used the PerturBench-compatible implementation released with the code.
- Analysis of regulatory-gene composition across generation steps.: A DeepSEM-derived control-cell GRN classified the 1,000 modeled genes into 303 regulatory genes with outgoing edges and 697 regulated genes.The analysis calculated regulatory- and regulated-gene proportions at each generation step.
- Comparison of ∆Order between perturbation-specific and non-specific TFs.: Perturbation-specific TFs were identified from DoRothEA regulon activity using perturbation-versus-control expression changes, while RPM-promoted genes were defined by ∆Order > 0 and compared with top differentially expressed genes.DoRothEA regulons were restricted to modeled TF-target pairs, and the top 100 genes by absolute differential-expression score defined each perturbation’s DE set.
- The Results of GSEA Analysis: RPM-promoted genes are strongly enriched in cell-cycle-related pathways, including cell-cycle and G2-M transition pathways, according to GSEA of averaged ∆Order values.GSEA used gseapy.prerank with Reactome Pathways 2024; positive NES denotes enrichment among genes promoted earlier by RPM.