Source-linked AI summary

scDEFT: A deep learning framework for drug-effect prediction and counterfactual reasoning

Murthy Devarakonda

arXiv:2609.10831v1q-bio.QMcs.LG

TL;DR

scDEFT addresses how the same drug can produce different patient trajectories, a question enabled by matched longitudinal single-cell profiles but not resolved by existing perturbation or compositional methods. It learns drug-conditioned representations with decoupled prediction heads and a backward attribution stage, achieving reproducible state-change prediction and responder stratification while supporting gene-program and counterfactual analyses.

  • Problem

    Existing single-cell perturbation methods often fail to beat linear baselines and are not designed to contrast responders with non-responders or attribute that contrast to interpretable genes.

  • Method

    scDEFT treats drug action as a transition operator, learns drug-conditioned cell representations with separate neighborhood-level heads, and traces predictive dimensions back to genes and programs.

  • Results

    Neighborhood aggregation raises state-change prediction to +0.273, or 45% of reproducibility headroom, while the model identifies responder-associated programs and supports counterfactual drug-effect prediction.

  • Takeaways & Limitations

    The framework provides predictive and explanatory outputs that nominate co-targets, including programs associated with unresolved treatment responses in non-responders.

  • Takeaways & Limitations

    Performance is bounded by the available data scale, comprising 51 donors, three cohorts, and two drug classes, and attribution results depend on analysis choices.

Abstract

from arXiv · show

Longitudinal single cell atlases now capture matched pre treatment and post treatment states from responders and non responders, presenting an opportunity to mechanistically explain why two patients on the same drug diverge. We introduce scDEFT (single cell Drug EFfect Transducer), which treats a drug as a conditioning operator on cell representations, enabling prediction and explanation. In scDEFT, feature wise linear modulation produces drug conditioned cell latents, learned under abundant per cell supervision and then frozen. Two independent heads aggregate those latents over shared transcriptional neighborhoods to predict drug induced state change and responder status. A backward stage ranks the latent dimensions by how strongly they separate responders from non responders and maps them to genes under a cell composition control. On a harmonized inflammatory bowel disease atlas of 1.16 million cells, three cohorts and two drug classes, scDEFT predicts state change at 45% of the baseline to reproducibility ceiling headroom and stratifies responders before treatment at AUROC 0.70, where standard predictors remain at chance. These predictions and the drivers behind them support target and co target nomination, patient stratification, and counterfactual prediction of unseen drug cohort effects.

Introduction

scDEFT addresses why patients receiving the same drug diverge by learning drug-conditioned cell representations that predict treatment-state changes and response, then tracing predictive dimensions to genes and programs.

  • Motivation: Matched pre- and post-treatment single-cell profiles with independent response labels make patient-specific treatment trajectories accessible in longitudinal cohorts.IBD provides such a setting, although roughly half of patients fail anti-TNF therapy and no single biomarker explains the difference.
  • Gap: Existing perturbation models often fail to beat linear baselines and do not contrast responders with non-responders or attribute differences to interpretable genes.Compositional approaches capture population entry or exit but not the underlying dynamics.
  • Approach: scDEFT learns drug-conditioned cell representations and uses separate neighborhood-level heads to predict state change and responder status.Its attribution stage ranks dimensions separating responders from non-responders and maps them to genes and programs.
  • Approach: The framework treats drug action as a transition operator and aggregates cells by transcriptional neighborhood to extract signal from noisy single-cell data.The resulting gene programs and co-targets are prioritized as hypotheses for experimental testing.

Results

scDEFT learns drug-conditioned representations from paired longitudinal single-cell data, then uses neighborhood-level prediction and backward attribution to model treatment response. Across IBD cohorts, neighborhood aggregation improves state-change and prospective responder prediction, while attribution identifies compartment-resolved programs and supports counterfactual and stratification analyses.

  • Framework and dataset: scDEFT learns drug-conditioned pre-treatment cell latents from per-cell state-change supervision, freezes them, and gives separate heads for neighborhood state change and response prediction.The atlas contains 1.16M cells from three longitudinal cohorts spanning anti-TNF and JAK-inhibitor treatments.
  • State-change prediction: Neighborhood aggregation lifts state-change prediction to +0.273, or 45% of reproducibility headroom, versus +0.082 for the per-cell model.This identifies transcriptional neighborhoods as the reproducible unit of donor-specific drug-effect signal.
  • Responder prediction: Prospective responder prediction reaches AUROC 0.70 [95% CI 0.56–0.84], while removing neighborhood aggregation yields 0.55 AUROC and leaves comparators’ intervals including chance.The prospective pipeline uses pre-treatment cells only; the retrospective pipeline reaches AUROC 0.93 from post-treatment cells.
  • Gene-driver attribution: Backward attribution recovers compartment-resolved responder programs, including epithelial barrier programs in responders and inflammatory TNF–NF-κB and MHC class-II programs in myeloid non-responders.Nine of ten myeloid drivers were classified as within-cell programs, with composition control reclassifying one S100A8/9⁺FCN1⁺ monocyte dimension as cell identity.
  • Applications: The myeloid inflammatory and antigen-presentation programs converge on JAK1/JAK2 as a proposed co-target in the interferon-γ → JAK–STAT1 → IRF1 → CIITA → MHC class II axis.The co-target is presented as a target hypothesis derived from the attributed programs.
  • Applications: Prospective stratification could support pre-treatment cohort enrichment, treatment selection, or rescue analyses, but confirmation outside the three cohorts remains future work.The paper presents these uses as supported applications rather than a claim to have solved IBD.
  • Applications: The tofacitinib operator predicts the observed tofacitinib shift more accurately than adalimumab on Thomas-UC cells, with agreement in 86% of 56 neighborhoods.The counterfactual comparison has cosine 0.289 versus 0.185 and one-sided Wilcoxon P = 1.7 × 10⁻⁹.

Discussion

scDEFT combines neighborhood-level prediction with backward attribution to make drug-effect modeling both predictive and explanatory. Its shared representation supports cross-cohort analysis, counterfactual drug queries, and actionable hypotheses, while current performance remains constrained by limited scale.

  • Discussion: The shared neighborhood design lifts state-change prediction to nearly half its reproducibility ceiling, while decoupling outperforms coupled objectives at realistic sample sizes.The framework uses neighborhoods that are fine enough to localize programs and coarse enough to remain reproducible.
  • Discussion: scDEFT is simultaneously predictive and explanatory, linking prospective non-responder programs and co-target nominations to the model’s predictive representation.The backward explanation is tied to the same representation that produces predictions.
  • Discussion: Compared with pseudo-bulk differential expression, scDEFT agrees on pro-repair macrophage and oxidative-metabolic axes but differs on MHC class-II and resident-macrophage signals.The analyses differ in whether they control within subtype composition.
  • Discussion: The atlas-scale framework uses a shared embedding and cross-cohort sign filtering so retained signals are common across independent studies rather than batch-specific.Its neighborhood assignment is linear in cell number and avoids an explicit graph over millions of cells.

Methods

The methods assemble a harmonized longitudinal IBD atlas and learn drug-conditioned cell representations that frozen downstream heads read over shared transcriptional neighborhoods. Separate state-change, response, attribution, and counterfactual procedures support standardized benchmarking and interpretation.

  • Atlas assembly: The atlas contains 1,156,405 cells from three longitudinal IBD cohorts, with 51 donors having paired pre/post samples and harmonized clinical metadata.The cohorts cover adalimumab, infliximab, and tofacitinib.
  • Neighborhood prediction: A shared K = 80 neighborhood scaffold defines the same transcriptional coordinates across cohorts, and gated attention pools frozen cell latents into neighborhood vectors for shift prediction.The pooled neighborhood vector has dimension 784, while the predicted shift has dimension 768; only attention and projection parameters are trained at this stage.
  • Shared representation: scDEFT uses 768-dimensional Geneformer-V2 cell embeddings, compartment context, and drug embeddings to produce drug-conditioned 784-dimensional latents through FiLM modulation.The drug generator supplies feature-wise scale γ(d) and shift β(d), while its output is initialized to leave an untrained drug unchanged.
  • Response and benchmarking: The prospective response head averages frozen latents across neighborhoods, reduces donor vectors to 30 principal components, and applies L2-regularized logistic regression under donor-grouped cross-validation.The benchmarking protocol compares state-change methods on a standardized scale and includes neighborhood and gene-space alternatives.
  • Backward attribution: Backward attribution ranks latent dimensions separating responders from non-responders, maps them to correlated genes after artifact masking, and applies composition controls before pathway testing.The analysis is performed by cohort and compartment on labeled pre-treatment cells.
  • Counterfactual queries: Counterfactual queries set the frozen model’s drug token to a queried drug and perform a forward pass without retraining, then read predicted neighborhood shifts from the state-change head.This permits predictions for drugs absent from the cohort.

Supplementary Note 1

scDEFT separates representation learning from downstream prediction because finite neighborhood-level supervision can undercut jointly trained multi-task representations. The winning design trains latents with abundant per-cell state signal, freezes them, and gives each task its own reader.

  • Decoupled versus end-to-end training: The decoupled configuration trains the drug-conditioned trunk and latents on per-cell state signal, then freezes the representation for separate neighborhood state and response heads.This design uses one representation with multiple decoupled readers.
  • Decoupled versus end-to-end training: Joint multi-task training consistently performed poorly, while symmetric coupling degraded state prediction and left response prediction near chance.An asymmetric end-to-end unfreeze also failed to beat freezing the representation.

Supplementary Note 2

SHAP and LIME attribute individual outputs to input features, whereas scDEFT requires a population-level contrast over frozen drug-conditioned representations to separate responders from non-responders.

  • Supplementary Note 2: SHAP and LIME are mismatched to scDEFT because they explain one model output for one instance rather than a responder-versus-non-responder population contrast.scDEFT compares representation elements across donor groups defined by response labels.

Supplementary Note 3

scDEFT identifies compartment-specific baseline programs separating inflammatory bowel disease responders from non-responders, broadly agreeing with prior cell-state findings while adding within-cell attribution. Its interpretation is constrained by disease pooling and analysis choices.

  • Epithelial compartment: Epithelial separation is dominated by responder-enriched antimicrobial and secretory barrier programs versus non-responder-enriched cell-cycle and mature absorptive-colonocyte programs.The cell-cycle and absorptive-colonocyte programs show especially strong E2F, G2–M, and oxidative-phosphorylation enrichment.
  • Myeloid compartment: The myeloid drivers separate future non-responders through inflammatory TNF–NF-κB and MHC class-II programs, while repair, regulatory, and oxidative-phosphorylation programs favor responders.The TNF–NF-κB driver is the largest single myeloid driver, and MHC class-II is the most reproducible result in that compartment.
  • Concordance and added analysis: The compartment-resolved drivers broadly recapitulate prior cohort associations while deriving within-cell programs from drug-conditioned latent attributions rather than abundance and differential-expression testing.The approach adds composition-controlled within-subtype gaps and ranks latent dimensions by their effects on the drug-conditioned representation.
  • Differences and unresolved findings: The analysis does not reproduce the prior interferon split or several T-cell and epithelial-damage associations, although it recovers the inflammatory-monocyte program after restricting correlations to baseline cells.MHC class-II shows a sharper compartment-specific split, loading toward responders in epithelium and non-responders in myeloid cells.
  • Caveats: Agreement is qualified because Crohn’s disease and ulcerative colitis are pooled and attribution outcomes depend on cell selection, ranking, exclusions, and composition control.Some comparator effects were disease-specific, including baseline epithelial frequency in Crohn’s disease.

Supplementary Note 4

scDEFT converts pre-treatment responder-associated programs into cell-type-specific target hypotheses. It also reveals upstream controls and distinguishes stratification markers from plausible therapeutic targets.

  • Targets and co-targets: Pre-treatment drivers are ranked by responder separation and linked to cell types and genes, yielding patient-grounded candidate targets and co-targets.Processes elevated only in future responders are framed as treatment gaps where new drugs or add-ons could be sought.
  • Upstream controls: The two top myeloid drivers converge on JAK1/JAK2, while CD40–CD40L and calprotectin-sensing blockade provide alternative treatment strategies.The inferred pathway is interferon-γ → JAK–STAT1 → IRF1 → CIITA → MHC class II.
  • Targeting cautions: The analysis cautions that suppressing antigen presentation could be broadly immunosuppressive because the axis reflects interferon tone and has opposite associations in epithelium and myeloid compartments.Pooled analyses would not expose this compartment-specific conflict.
  • Stratification: Epithelial drivers are positioned primarily for patient stratification, with routine MKI67 or TOP2A staining proposed to identify non-responders.Responder-enriched DUOX2, LCN2, PLA2G2A, and PIGR are framed as protective programs to preserve.

Supplementary Note 5

Compared with differential gene expression, scDEFT recovers similar responder-associated myeloid biology while controlling composition and linking programs directly to predictive response models.

  • Concordant biology: Both analyses identify pro-repair macrophage and oxidative-metabolic programs on the responder side, despite using different readouts.Differential expression emphasizes CD163, FABP4, HMOX1, and MT1M, whereas latent attribution identifies the TREM2/DAP12 lysosomal axis.
  • Comparison with differential expression: scDEFT reaches the same headline biology as manually decontaminated differential expression while automatically filtering composition- and cohort-specific noise.Its within-subtype composition control retains local differences, and cross-cohort sign filtering removes technical and cohort-specific signals.

Supplementary Note 6

scDEFT addresses a clinical-response question that cell-line perturbation models cannot answer alone. Its held-out treatment-shift accuracy is comparable to a foundation-model pipeline, but in longitudinal primary patient tissue.

  • Clinical-response setting: Cell-line foundation models learn molecular perturbation effects from cross-sectional in-vitro data but lack matched patient trajectories and clinical response labels needed to predict who responds.scDEFT instead trains on longitudinal primary patient data, tying predictions and nominated genes to responders and non-responders.
  • Foundation-model context: Tahoe-100M comprises over 100 million cells from 50 cancer cell lines exposed to more than 1,100 small molecules, illustrating the scale and in-vitro scope of the foundation-model training data.These models predict molecular tasks such as transcriptional effects, gene essentiality, or cell identity.
  • Benchmark comparison: In held-out cellular-context or donor prediction, a Tahoe-x1 transition-model pipeline is close to scDEFT accuracy, while scDEFT operates on primary tissue, real therapies, and matched pre- and post-treatment samples.The comparison is presented as complementary rather than competitive.
Loading 2609.10831v1…