Source-linked AI summary

Path2ST: Hierarchical Cell-Tissue Grounded Cross-Modal Translation for Spatial Transcriptomics

Ruochen Liu, Wei Lou

arXiv:2608.14710v1cs.CVcs.AIcs.CL

TL;DR

Spatial transcriptomics is informative but costly, while H&E images are widely available yet do not directly measure molecular profiles. Path2ST formulates H&E-to-ST prediction as hierarchical cross-modal semantic translation, combining biological conditioning, coarse-to-fine autoregressive generation, and SpectraLoss; experiments report state-of-the-art performance across three benchmarks.

  • Problem

    ST is costly and labor-intensive, motivating inference of spatial gene expression from widely available H&E images whose morphology may encode latent molecular signals.

  • Method

    Path2ST jointly encodes cellular composition, cell-type prototypes, and tissue context, then generates expression autoregressively from global profiles to gene-group detail.

  • Results

    Path2ST achieves state-of-the-art performance across three public benchmarks covering diverse species and tissue types.

  • Takeaways & Limitations

    The framework provides biologically grounded and statistically meaningful transcriptomic synthesis with reported biological interpretability across multiple benchmarks.

  • Takeaways & Limitations

    The formulation assumes conditional independence across spots while generating each spot’s gene expression profile.

Abstract

from arXiv · show

Predicting spatial gene expression from hematoxylin and eosin (H\&E)-stained images offers a cost-effective alternative to spatial transcriptomics (ST). However, existing methods treat H\&E images as generic visual inputs and ignore their intrinsic biological hierarchy, where spatially organized cell types collectively form functional tissue microenvironments that govern local gene expression programs. To bridge this gap, we formulate H\&E-to-ST prediction as a cross-modal semantic translation task and propose Path2ST, a hierarchically grounded autoregressive framework featuring three key components: (i) a Hierarchical Cell-Tissue Conditioning mechanism that fuses explicit and implicit cellular features with tissue-level semantic representations to construct hierarchical conditioning signals; (ii) a Scale-Adaptive Autoregressive Generation process over a hierarchical semantic vocabulary, enabling coarse-to-fine, biologically consistent expression synthesis; and (iii) SpectraLoss, a full-spectrum objective that jointly enforces ordinal fidelity, models transcriptional bursts, and aligns semantic structures with cell types. Extensive experiments on three datasets demonstrate state-of-the-art performance, validating that Path2ST generates highly accurate and spatially coherent transcriptomic profiles. The related code is released at https://github.com/RuochenLiu23/Path2ST.

1 Introduction

The paper frames H&E-to-ST prediction as hierarchical cross-modal semantic translation because gene expression reflects cellular composition, cell states, and tissue microenvironment. Path2ST addresses this structure through biologically grounded conditioning, coarse-to-fine autoregressive generation, and full-spectrum supervision, achieving state-of-the-art results on three benchmarks.

  • Motivation: ST provides genome-wide expression with spatial context, but its clinical and large-scale use is limited by cost, instrumentation, and labor-intensive workflows.H&E slides are inexpensive, widely available, and embedded in standard pathology pipelines, motivating image-based molecular inference.
  • Research gap: Existing methods often treat histology as generic visual input, overlooking how cellular composition, intrinsic cell states, and tissue microenvironment jointly shape spot-level expression.They also commonly represent gene expression as a flat high-dimensional target rather than a structured transcriptional program.
  • Path2ST framework: Path2ST fuses explicit cell-type composition, implicit cell-type prototypes, and tissue-context features into a unified hierarchical conditioning signal.Adaptive gating modulates the feature streams to align visual representations with biologically meaningful cellular identities.
  • Path2ST framework: Its scale-adaptive autoregressive decoder generates spot expression from global profiles toward gene-group-level detail using a hierarchical semantic vocabulary structured by gene co-expression.Dynamic remodulation of semantic conditioning preserves coherence across granularity levels.
  • Path2ST framework: SpectraLoss jointly constrains predictive fidelity, transcriptional distribution statistics, and cell-type semantic structure during expression generation.The objective is designed to match observed values while respecting statistical and semantic properties of spatial transcriptomic data.
  • Results: Path2ST achieves state-of-the-art performance across three public benchmarks spanning diverse species and tissue types.The supplied introduction reports this as the principal experimental outcome without providing a numerical headline metric.

2 Related Work

Prior methods infer spatial expression through regression, contrastive alignment, or generative modeling, but do not explicitly combine cell–tissue semantic hierarchy with structured gene generation. Path2ST instead uses unified biological conditioning and hierarchical autoregressive decoding for direct transcriptional generation.

  • Regression-Based Methods: Regression-based methods learn mappings from image features to spot-level gene expression, using architectures such as DenseNet and Vision Transformers.Examples include ST-Net, DeepSpaCE, HisToGene, and Hist2ST.
  • Contrastive Learning-Based Methods: Contrastive methods align histopathology and transcriptomic signals in a shared latent space and may predict expression through nearest-neighbor retrieval.BLEEP uses bimodal embeddings and retrieval, while NH22ST combines dual-scale contrastive learning with hypergraph modeling.
  • Path2ST: Path2ST differs by explicitly combining unified biological conditioning with hierarchical autoregressive decoding for direct transcriptional generation.This design targets cell–tissue semantic hierarchy and structured gene generation together.
  • Generative Methods: Generative methods model expression uncertainty or joint distributions using conditional diffusion, flow matching, and next-scale autoregressive generation.STEM, STFlow, and GenAR represent these generative directions.

3 Methodology

Path2ST formulates H&E-to-ST prediction as conditional autoregressive cross-modal translation, using tissue context and cell-type composition to generate spatial gene expression. Its methodology combines hierarchical conditioning, coarse-to-fine generation, and SpectraLoss supervision.

  • Problem formulation: Each spot is represented by a gene-count vector, and the aggregate matrix characterizes the tissue’s global molecular landscape.Spot-level profiles record captured mRNA counts across the gene set.
  • Problem formulation: The task maps whole-slide histology images to spot-level spatial transcriptomic profiles through a conditional autoregressive process.The WSI is treated as the source modality and the spatial transcriptomic matrix as the target.
  • Hierarchical Cell-Tissue Conditioning: Hierarchical conditioning combines spatially aware tissue embeddings with cell-type proportions derived from segmented cells within each spot.Cell bags use cells within a radius of 112 pixels, while tissue patches are encoded by a pretrained pathology model.
  • Hierarchical Cell-Tissue Conditioning: Dynamic gating adaptively modulates how strongly fused cellular representations are injected into tissue features.The gate is input-dependent and uses cell-composition and tissue-context projections for dimension-wise modulation.
  • Hierarchical Cell-Tissue Conditioning: Hybrid explicit–implicit cross-attention uses proportion-weighted cell-type prototypes and tissue-conditioned queries to produce cellular representations.Explicit keys and values encode cell-type composition, while queries retrieve relevant latent cellular semantics under tissue context.
  • Scale-Adaptive Autoregressive Generation and SpectraLoss: Hierarchical generation decomposes expression synthesis into coarse-to-fine scales, while SpectraLoss jointly enforces predictive, distributional, and semantic constraints.The loss combines adaptive Gaussian-target KL divergence, zero-inflated negative binomial likelihood, and soft-positive contrastive learning.

4 Experiments

Experiments use three spatial transcriptomics datasets spanning prostate cancer, breast cancer, and healthy mouse brain, with paired histology and transcriptomic data.

  • Datasets: The evaluation covers PRAD, HER2ST, and Healthy Mouse Brain spatial transcriptomics datasets.PRAD contains paired prostate cancer ST and histology data; HER2ST contains paired breast cancer data; Healthy Mouse Brain covers mouse brain tissue.
  • Datasets: PRAD includes benign, transitional, and tumor regions across multiple Gleason grades.The dataset uses 55 µm spots and contains 1,418–4,079 spots per slide.
  • Datasets: HER2ST contains normal, immune-infiltrated, in situ, and invasive carcinoma regions across 13,594 spots.Its paired histology images use 20× magnification and 100 µm spots.

4.2 Implementation Details

Implementation uses fixed 224-pixel patches, multi-scale autoregressive generation, raw-count prediction, and log2-transformed predictions during evaluation.

  • Implementation details: All patch sizes are set to 224 pixels, and generation scales are configured as (1, 4, 8, 40, 100, 200).Experiments run on an NVIDIA A40 GPU.
  • Implementation details: The model directly predicts raw counts, with log2 transformation applied to predictions at evaluation.
  • Implementation details: MEND145, SPA148, and NCBI667 serve as test sets for PRAD, HER2ST, and Healthy Mouse Brain, respectively.

4.3 Evaluation Metrics

Evaluation uses correlation and absolute or squared numerical error metrics, with PCC additionally summarized over top-ranked genes.

  • Metrics: PCC measures correlation between predicted and true expression values for each gene across all spots.
  • Metrics: PCC-10, PCC-50, and PCC-200 average PCC over the top 10, 50, and 200 genes ranked by PCC.
  • Metrics: MSE and MAE quantify numerical errors in predicted expression values.

4.4 Comparison with Existing Methods

Path2ST achieves the best overall performance across three benchmarks, outperforming GenAR on cancer and healthy-brain datasets across multiple metrics.

  • Cross-dataset comparison: Path2ST achieves the best overall performance across PRAD, HER2ST, and Healthy Mouse Brain.The comparison includes contrastive, regression, and generative baselines, including GenAR.
  • PRAD: 6.5% and 6.6% gains over GenAR on PRAD improve PCC-10 and PCC-200, while MSE and MAE decrease by 0.186 and 0.046.
  • HER2ST: On HER2ST, Path2ST improves over GenAR by 1.2% on PCC-10 and 1.7% on PCC-50.
  • Healthy Mouse Brain: On Healthy Mouse Brain, Path2ST surpasses GenAR by 3.7% on PCC-10 and 4.7% on PCC-200.The result is reported under limited-sample conditions and on a different species and tissue type.

4.5 Ablation Study

The ablation study shows that hierarchical cellular conditioning and semantic contrastive alignment provide the largest reported gains, while autoregressive generation and count modeling add further improvements.

  • Hierarchical Cell-Tissue Conditioning improves PCC-200 from 0.519 to 0.549 and reduces MSE from 1.148 to 1.080.The gain supports incorporating cellular context into autoregressive conditioning.
  • Scale-Adaptive Autoregressive Generation yields consistent improvements across all reported metrics.
  • Adding ZINB Loss further improves count modeling, increasing correlation and reducing error metrics.
  • The complete configuration achieves PCC-10 of 0.767 and MSE of 1.005 after adding Soft-Positive Semantic Contrastive Loss.The largest improvements come from hierarchical cellular conditioning and semantic contrastive alignment.

4.6 Hyperparameter Analysis

The hyperparameter analysis examines the number of learnable tissue-context queries in the asymmetric conditioning design.

  • The module uses explicit cell-type proportions as Keys and Values and tissue-context-guided learnable Queries.The Queries retrieve expression-relevant combinations from the cell-type semantic prototype space.
  • The PRAD analysis studies how the query number N_q affects this conditioning design.

4.7 Visualization and Explainability

Visualization results show close agreement between predicted and ground-truth spatial expression patterns and connect selected cancer biomarkers with neoplastic-cell distributions.

  • Predicted spatial expression patterns for SORD, TSPAN1, and TMPRSS2 closely match the ground truth.
  • Cell-type classification is visualized across a whole-slide image containing over 60,000 cells.
  • SORD, TSPAN1, and TMPRSS2 are prostate cancer biomarkers associated with disease progression and highly enriched in neoplastic cells.

4.8 Conclusion

The conclusion presents Path2ST as a hierarchical cross-modal translation framework that jointly models cellular composition, tissue context, and gene-expression structure. Experiments report state-of-the-art performance with biological interpretability across multiple benchmarks.

  • Path2ST models pathological images through hierarchical cell-tissue semantic translation for spatial transcriptomic synthesis.
  • Its adaptive explicit-implicit mechanism aligns cellular composition with the tissue microenvironment across levels.
  • Scale-adaptive autoregressive generation models gene co-expression relationships while maintaining semantic consistency across scales.
  • SpectraLoss supervises biological statistical properties, semantic consistency, and numerical fidelity.These dimensions are described as making generated profiles biologically and statistically meaningful.
  • Experiments demonstrate state-of-the-art performance across multiple benchmarks with strong biological interpretability.The conclusion highlights potential for cost-effective transcriptomic synthesis in digital pathology.
Loading 2608.14710v1…