Source-linked AI summary
DeMixPert: Decomposed Response Modeling with Gaussian Mixtures for OOD Single-Cell Perturbation Prediction
Jiawen Liu, Xuechenxiao Cao, Yutong Li, Bing Liu, Jiaming Liang, Tinghe Zhang, Xiaoqi Sheng, Hongmin Cai
TL;DR
DeMixPert addresses OOD single-cell perturbation prediction by separating basal-state-dependent, perturbation-specific, and population-level response components. It combines adaptive Gaussian prototypes with an invertible network and reports improved recovery of perturbation-specific responses and population distributions across unseen settings.
Problem
Predicting unseen perturbation responses requires modeling perturbation-specific transcriptional effects and heterogeneous cellular responses beyond average expression shifts.
Method
DeMixPert decomposes responses into systematic, perturbation-specific, and population-level components, modeling the latter with context-adaptive Gaussian prototypes and an invertible network.
Results
DeMixPert improves OOD perturbation prediction by recovering perturbation-specific responses while preserving population-level distributional fidelity.
Takeaways & Limitations
The framework captures heterogeneous perturbation responses and condition-specific population structures under unseen-perturbation settings.
Takeaways & Limitations
The method assumes control and perturbed populations are unpaired and randomly samples a control cell as each perturbed cell’s reference.
Abstract
from arXiv · showhide
Predicting transcriptome-wide responses to unseen genetic perturbations remains a major computational challenge because accurate prediction requires recovering both perturbation-specific transcriptional shifts and heterogeneous cellular responses. Existing methods often entangle deterministic response structure with stochastic population-level variation, causing dominant shared patterns to mask weaker perturbation-specific signals and impair distributional modeling. To address these challenges, we propose \textbf{DeMixPert}, an approach for Decomposed response Modeling with Gaussian Mixtures for Out-Of-Distribution (OOD) single-cell Perturbation prediction. DeMixPert decomposes perturbation-induced changes into a basal-state-dependent systematic response, a perturbation-specific response, and population-level variation. The systematic component is derived from the basal state encoded from control-cell expression, whereas the perturbation-specific component is inferred from pretrained target embeddings for unseen-target generalization. DeMixPert models population-level variation using a Gaussian prototype Invertible Network and adaptively combines reusable Gaussian prototypes according to the basal state and perturbation condition. The resulting mixture is mapped to a condition-specific variation distribution. Sampled variations are integrated with the systematic and perturbation-specific components, followed by joint decoding with the basal state to reconstruct perturbed-cell gene expression. Experimental results show that DeMixPert effectively captures heterogeneous single-cell perturbation responses and achieves superior performance across unseen-perturbation settings. The source code is made publicly available upon publication.
Introduction
Single-cell perturbation profiling enables cell-resolved transcriptomic analysis but cannot feasibly measure all combinatorial conditions. DeMixPert addresses OOD prediction by separating systematic, perturbation-specific, and population-level response components.
- Single-cell profiling links controlled genetic interventions to transcriptome-wide, cell-resolved measurements of regulatory responses.
- Experimental scalability is limited by the combinatorial growth of perturbation targets, cellular states, and biological contexts.
- OOD prediction must recover both basal-state-dependent systematic effects and perturbation-specific transcriptional responses.
- DeMixPert separates systematic response, perturbation-specific response, and population variation into dedicated modules.
- The framework adaptively combines reusable Gaussian prototypes using basal-state and perturbation representations to model condition-specific variation distributions.
- Quantitative analyses report improved OOD prediction by recovering perturbation-specific responses while preserving population-level distributional fidelity.
Related Work
Related methods address perturbation prediction through latent transformations, biological knowledge, and population-distribution modeling. These approaches differ in how they represent perturbation effects, generalize to unseen targets, and capture heterogeneity.
- Latent-variable methods represent responses through state transformations or factorized perturbation components.
- Knowledge-guided methods use gene representations, biological priors, gene-relation graphs, or external embeddings for unseen-target generalization.
- Population-distribution methods model control-to-perturbed transitions using cell sets, distributional objectives, flow matching, or conditional invertible networks.
- Gaussian mixtures provide flexible population representations, while normalizing flows transform tractable source distributions through invertible mappings.
Methodology
DeMixPert constructs response predictions from unpaired control and perturbed cells by decomposing latent responses and modeling residual heterogeneity with adaptive Gaussian mixtures and invertible transformations.
- Preliminaries: DeMixPert decomposes each predicted response into systematic, perturbation-specific, and population-level variation components.
- Preliminaries: The method uses randomly sampled control cells as references because control and perturbed populations are unpaired.
- Deterministic Response Center Learning: Basal-state and perturbation representations define separate systematic and perturbation-specific responses, whose sum forms the deterministic response center.
- Deterministic Response Center Learning: Pretrained scGPT target embeddings provide the perturbation-specific response representation for OOD generalization.
- Gaussian prototype Invertible Network: A Gaussian prototype Invertible Network models residual population variation by adapting shared Gaussian prototypes to basal-state and perturbation context.
- Gaussian prototype Invertible Network: Conditional mixture samples pass through bijective coupling layers to generate variation samples, which are added to the deterministic center.
- Expression Reconstruction: The response latent and basal-state representation are concatenated and decoded to reconstruct post-perturbation gene expression.
Experiments
DeMixPert is evaluated on unseen single-gene and combinatorial perturbations using distributional, expression, and perturbation-specific metrics, alongside ablations, sensitivity analyses, and interpretability studies. Results indicate strong OOD recovery, while decomposition and adaptive Gaussian-mixture modeling contribute complementary benefits.
- Quantitative comparison: DeMixPert is benchmarked on four datasets covering unseen single-gene and combinatorial perturbations, using top-100-DEG and all-gene metrics.The evaluation compares DeMixPert with GEARS, scGPT, GenePert, scDFM, and STATE under condition-disjoint splits.
- Quantitative comparison: 22.000 and 10.733 C-DEGs, 0.904 and 0.947 PDS, and 0.171 and 0.448 E-Dist are reported for Papalexi and Adamson, respectively.DeMixPert substantially improves perturbation-specific and distributional recovery on these unseen single-gene datasets.
- Quantitative comparison: 0.981 and 0.938 PDS and 0.374 and 0.155 E-Dist are achieved on Norman and Replogle, respectively, for unseen combinatorial perturbations.DeMixPert reaches 31.650 C-DEGs on Norman and 12.378 on Replogle.
- Ablation study: Removing decomposition lowers Centroid Acc/PDS to 0.240/0.244 and 0.733/0.736 on Adamson/Replogle, while removing the gene effect reduces C-DEGs to 8.773/9.733.Removing the Gaussian-mixture module raises C-DEGs but worsens all other metrics, indicating that the deterministic center alone does not recover population distributions.
- Sensitivity analysis: The reference setting K = 8 and λgm = 0.01 achieves the best balance across DEG recovery, distribution modeling, expression reconstruction, and perturbation identification.Larger K or higher λgm can improve perturbation identification but affect other objectives, revealing a fidelity–separability trade-off.
- Interpretability analysis: Removing the systematic component raises cluster purity from 0.308 to 0.769 on Replogle and from 0.555 to 0.825 on Adamson.The comparison evaluates whether basal-state-dependent variation interferes with perturbation-associated structure in predicted response space.
- Interpretability analysis: Prototype-weight divergence is positively associated with population-distribution distance, with ρ = 0.462 on Adamson and ρ = 0.512 on Replogle, both with Mantel p < 0.001.Dα is the JS divergence between prototype-weight vectors, while Dϵ is the energy distance between population distributions.
Conclusion
DeMixPert decomposes OOD single-cell perturbation responses into systematic, perturbation-specific, and population-level components. Its adaptive Gaussian prototype composition models condition-specific cellular heterogeneity and supports recovery of perturbation-specific responses.
- Conclusion: DeMixPert disentangles systematic, perturbation-specific, and population-level components through dedicated modeling mechanisms.A Gaussian prototype Invertible Network adaptively composes prototypes to characterize cellular heterogeneity.
- Conclusion: Extensive experiments demonstrate accurate recovery of perturbation-specific responses under OOD settings.The framework also captures condition-specific population structures, providing a structured interpretation of heterogeneous cellular responses.
Perturbation Embedding Construction
DeMixPert constructs perturbation embeddings from pretrained scGPT gene representations. Single-gene conditions use one target embedding, while combinatorial conditions average target embeddings equally.
- Perturbation embeddings are constructed from pretrained scGPT gene representations.
- Single-gene perturbations use the embedding of their targeted gene.
- Combinatorial perturbations use an equal-weight average of their target-gene embeddings.
- Conditions with targets unavailable in the frozen embedding table are removed before split generation.
Default Model Configuration
The default DeMixPert architecture and optimization settings are shared across datasets, while dataset-specific training settings are reported separately in the supplementary appendix.
- The default architecture and optimization configuration is shared across all datasets.Dataset-specific training settings are provided in Supplementary Appendix H.
- The model uses a 256-dimensional state latent space and a 128-dimensional response latent space.
- The configuration specifies 8 Gaussian components and 4 coupling layers.
- Optimization uses AdamW with a learning rate of 5 × 10^-5, global-norm gradient clipping at 5.0, and no learning-rate scheduler.The prototype temperature is 1.0.
- Dataset-specific batch sizes, evaluation intervals, Gaussian-mixture refresh settings, and training schedules are listed in Supplementary Appendix H.
Dataset Descriptions
DeMixPert is evaluated on four single-cell genetic perturbation datasets spanning single-gene and combinatorial settings, multiple perturbation technologies, and different cellular contexts.
- The evaluation covers Adamson, Papalexi, Norman, and Replogle.Adamson and Papalexi are single-gene datasets, while Norman and Replogle include combinatorial perturbations.
- Adamson uses Perturb-seq in K562 cells to measure responses to single-gene perturbations related to the unfolded protein response.
- Papalexi uses ECCITE-seq in THP-1 cells to measure CRISPR perturbations targeting immuneregulatory genes.Its conditions are single-gene perturbations.
- Norman contains CRISPR activation measurements in K562 cells involving single-gene and combinatorial perturbations.It studies transcriptional responses and genetic interactions.
- Replogle contains single-gene and dual-gene perturbations measured using direct guide-RNA capture technology.
Data Preprocessing
The preprocessing pipeline defines the response from matched control and perturbed cells and uses standardized frozen expression inputs with perturbation embeddings constructed separately.
- Data preprocessing covers quality control, normalization, gene selection, perturbation embedding construction, and response-reference definition.
- All models and metrics operate on the frozen logNor expression layer.
- The feature space combines 2,048 expression-variability-selected genes with available perturbation target genes.Gene selection occurs before condition-level splitting and is shared across five seeds and compared methods.
- Perturbation embeddings are constructed according to the procedure defined in Supplementary Appendix A.
- For each perturbed cell, a control cell is randomly sampled from the control population to define the response reference.
Out-of-Distribution Data Splitting Protocol
The evaluation uses perturbation-level splits in which non-control perturbation conditions are mutually disjoint across training, validation, and test sets. Control cells remain a shared reference for constructing perturbation responses and evaluating differential expression.
- Non-control perturbation conditions are mutually disjoint across the training, validation, and test sets.Control cells are not prediction targets in the split.
- Control cells are retained as a shared reference population for constructing perturbation responses and evaluating differential expression.
- The split therefore evaluates generalization to unseen perturbation conditions rather than memorization of cells from observed perturbations.
Single-Gene Perturbation Datasets
The experiments define dataset-specific perturbation splits, Gaussian-mixture fitting procedures, evaluation metrics, and ablations to assess DeMixPert’s response decomposition and population modeling. Results include improved separability after removing basal-state-driven variation and significant associations between prototype-weight profiles and residual population differences.
- Data splitting: Adamson and Papalexi non-control conditions are sorted, randomly permuted with a seed-specific generator, and divided into training, validation, and test sets.For datasets with at least three non-control conditions, validation and test sets are non-empty; split conditions are stored lexicographically.
- Data splitting: Norman and Replogle place all single-gene conditions in training and randomly divide dual-gene conditions between training, validation, and test.Training receives round(0.5Ndual) dual-gene conditions, while an odd remainder gives test one more condition than validation.
- Evaluation design: Five seeds, S = {17, 23, 29, 31, 37}, are used for every method, with each generated split saved before training and reused across comparisons.Condition-level separation targets generalization to unseen perturbations rather than memorization of observed perturbation cells.
- Ablation design: The ablations remove decomposition, the embedding-derived target-specific branch, the basal-state-dependent systematic branch, or population variation to isolate module contributions.Each variant is independently initialized and retrained on Adamson and Replogle rather than generated by post-hoc masking.