Source-linked AI summary

HistoAtlas: A Pan-Cancer Morphology Atlas Linking Histomics to Molecular Programs and Clinical Outcomes

Pierre-Antoine Bannier

arXiv:2603.16587v1q-bio.QMcs.CVeess.IV

TL;DR

HistoAtlas addresses the lack of a quantitative pan-cancer atlas for histopathology that links morphology with molecular biology and clinical outcomes. It systematically connects interpretable features from routine H&E slides to these data, recovering canonical programs and revealing compartment-specific immune associations and morphological subgroups with divergent outcomes.

  • Problem

    Histopathology lacks a quantitative pan-cancer atlas linking routinely generated H&E morphology with molecular biology and clinical outcomes.

  • Method

    HistoAtlas extracts 38 quantitative histomic features from 6,745 diagnostic slides across 21 TCGA cancer types and systematically links them to survival, molecular data, and immune subtypes.

  • Results

    The atlas links morphology to canonical molecular programs and clinical outcomes while identifying compartment-specific immune associations and morphologically distinct subgroups with divergent survival.

  • Takeaways & Limitations

    HistoAtlas provides an openly queryable, spatially traceable resource for systematic biomarker discovery from routine H&E slides.

  • Takeaways & Limitations

    Most survival associations have insufficient evidence because smaller cohorts provide limited statistical power, particularly for cancer types with fewer than 100 samples.

Abstract

from arXiv · show

We present HistoAtlas, a pan-cancer computational atlas that extracts 38 interpretable histomic features from 6,745 diagnostic H&E slides across 21 TCGA cancer types and systematically links every feature to survival, gene expression, somatic mutations, and immune subtypes. All associations are covariate-adjusted, multiple-testing corrected, and classified into evidence-strength tiers. The atlas recovers known biology, from immune infiltration and prognosis to proliferation and kinase signaling, while uncovering compartment-specific immune signals and morphological subtypes with divergent outcomes. Every result is spatially traceable to tissue compartments and individual cells, statistically calibrated, and openly queryable. HistoAtlas enables systematic, large-scale biomarker discovery from routine H&E without specialized staining or sequencing. Data and an interactive web atlas are freely available at https://histoatlas.com .

1 Introduction

HistoAtlas addresses the lack of quantitative, pan-cancer resources linking H&E morphology with molecular biology and clinical outcomes. It provides an interpretable, spatially traceable atlas that systematically evaluates histomic features across TCGA cancer types and biological domains.

  • Motivation: H&E slides encode quantitative cellular, nuclear, immune, and stromal information, but existing pan-cancer resources largely categorize or discard these morphological signals.Genomic, transcriptomic, proteomic, and epigenomic resources do not provide equivalent quantitative morphology.
  • Motivation: Existing resources and approaches do not combine interpretable morphology with systematic molecular linkage at pan-cancer scale or provide compartment-specific resolution.Prior deep-learning TIL mapping reported a single bulk density score without compartment-specific resolution or gene-expression linkage.
  • HistoAtlas: 38 quantitative histomic features were extracted from 6,745 TCGA diagnostic slides across 21 cancer types, with associations tested against survival, gene expression, mutations, copy-number variation, and immune subtypes.The atlas includes a pooled pan-cancer analysis and applies explicit correction families with evidence-strength badges.
  • HistoAtlas: Every association is spatially traceable to tissue compartments and individual cells, enabling compartment-resolved immune analysis beyond bulk H&E-derived TIL scoring.The study reports a stronger protective observational association between intratumoral lymphocyte density and survival than stromal lymphocyte density.

2 Results

HistoAtlas combined 6,745 diagnostic H&E slides across 21 TCGA cancer types with automated tissue and cell segmentation to quantify interpretable morphology. These features and morphological clusters showed biologically coherent links to molecular programs, immune biology, mutations, and survival, while analyses quantified batch effects and evidence limitations.

  • Atlas construction: HistoAtlas comprised 6,745 H&E-stained diagnostic slides spanning 21 TCGA solid-tumor cancer types.Twelve additional cancer types were excluded because their dominant cell morphologies fell outside the segmentation models’ training domain.
  • Atlas construction: Automated segmentation classified approximately 1.4 m^2 of tissue into five compartments, with tumor and stroma together covering over 90% of analyzed area.A second model identified more than 4.4 billion individual cells.
  • Molecular and immune associations: Morphological clusters aligned with canonical molecular programs, including immune rejection in Cluster 4 and Wnt/β-catenin signaling in Cluster 6.Cluster 4 was 76% THYM and showed immune rejection enrichment (δ = 0.67, Padj = 3.3 × 10^-40); Cluster 6 was 61% COAD and READ.
  • Prognostic morphology: Intratumoral lymphocyte density predicted favorable survival more strongly than stromal lymphocyte density across cancers.The pan-cancer hazard ratios were HR = 0.87 for intratumoral density and HR = 0.89 for stromal density.
  • Molecular and immune associations: In BRCA, morphology-to-molecular links included intratumoral lymphocyte density with CD8A (ρ = 0.59) and TIGIT (ρ = 0.63), and mitotic index with PLK1 (ρ = 0.56).The associations were multiple-testing corrected and based on n = 958 BRCA samples for the reported examples.
  • Cluster-level outcomes: Cluster-level survival analysis identified favorable survival for quiescent Cluster 2 (HR = 0.54) and adverse survival for hormone-driven Cluster 8 (HR = 1.37).Cluster 2 also had suppressed E2F-target proliferation scores, whereas the remaining clusters were not significant after BH correction.

3 Discussion

HistoAtlas links spatially resolved H&E-derived features to canonical molecular programs and clinical outcomes, with most audited claims supported or biologically plausible. The discussion highlights compartment-specific immune prognostic signals, transparency and reuse, resource gaps, and limitations requiring validation and extension.

  • Biological significance: HistoAtlas’s 38 spatially resolved histomic features recapitulate proliferation, kinase, EMT, and immune programs while stratifying clinical outcomes across cancer types.The discussion emphasizes comprehensive, statistically transparent links to survival, gene expression, and mutations.
  • Biological significance: 42 of 60 audited claims (70%) were established or literature-supported, while 12 (20%) were novel but biologically plausible.Five claims (8%) had uncertain mechanisms, and one apparent contradiction was resolved as spatial composition heterogeneity rather than genetic clonal diversity.
  • Biological significance: Intratumoral lymphocyte density showed strong prognostic protection, whereas stromal lymphocyte density showed weak protection, supporting the importance of immune localization.Prior bulk TIL scoring did not distinguish intratumoral from stromal compartments.
  • Context and resources: Compared with cBioPortal, the Human Protein Atlas, and The Cancer Imaging Archive, HistoAtlas adds continuous H&E morphometrics and statistical molecular and clinical linkages.cBioPortal lacks morphological features, the Human Protein Atlas focuses on semi-quantitative immunohistochemistry and single-protein correlations, and The Cancer Imaging Archive lacks a statistical layer.
  • Limitations: The atlas is limited by its 6,745-slide retrospective TCGA convenience cohort, coverage of 21 of 33 cancer types, and segmentation domains excluding 12 tumor types.The excluded types contain dominant populations outside the segmentation model’s training domain; interpretable features were also not benchmarked against foundation-model embeddings.
  • Future directions and reuse: Spatial transcriptomics validation, foundation-model integration, and extension beyond TCGA are proposed to strengthen biological validation and interpretability–predictive-power comparisons.The atlas also provides evidence-strength badges, minimum detectable effects at 80% power, and bidirectional spatial traceability from statistics to compartments and cells.

4 Methods · 4.1 Data acquisition · 4.2 Feature extraction

HistoAtlas acquired quality-controlled TCGA diagnostic H&E whole-slide images with matched clinical and molecular data, then extracted 38 histomic features using tissue and cell segmentation. Spatial features were standardized across compartments, while unreliable normal-epithelium features were excluded and ratio artifacts mitigated.

  • 4.1 Data acquisition: Diagnostic FFPE H&E whole-slide images from 21 TCGA solid-tumor cancer types were obtained through the GDC portal and quality-filtered for tissue area, artifacts, and clinical metadata.Slides were excluded when viable tissue was below 1 mm2, pen marks covered >20% of tissue, out-of-focus regions occurred, or vital status or follow-up time was missing.
  • 4.1 Data acquisition: Matched survival, demographic, stage, tissue-source, purity, expression, mutation, copy-number, immune-fraction, and immune-subtype data were retrieved from TCGA resources and matched by case barcode.Clinical endpoints included overall, disease-specific, disease-free, and progression-free survival; immune subtypes were C1–C6.
  • 4.2 Feature extraction: The atlas computed 38 quantitative histomic features spanning tissue composition, cell densities, nuclear morphology and kinetics, spatial organization, and spatial heterogeneity.The categories contained 3, 6, 8, 18, and 3 features, respectively.
  • 4.2 Feature extraction: Tissue segmentation used a CellViT-inspired model with a Phikon self-supervised ViT-B backbone, inference at 0.5 µm/px, 224×224-pixel tiles, and overlap-region majority voting.The model classified tiles into nine tissue classes, although the supplied passage truncates the class list.
  • 4.2 Feature extraction: HistoPLUS detected and classified nine morphological cell types, achieving mean panoptic quality [PQ] = 0.509, with per-class PQ ranging from 0.28 to 0.73.Performance was lowest for rare cell types such as eosinophils and apoptotic bodies.
  • 4.2 Feature extraction: Cell tiles used a 64-pixel (16 µm) overlap margin, and instances with centroids within 10 µm were deduplicated using union-find.The model was trained on annotations from eight cancer types and applied out-of-distribution to the remaining 13.
  • 4.2 Feature extraction: Spatial features used compartment masks at r = 8 µm/px, removed components below Amin = 2,048 µm2, and defined bands with signed Euclidean distance thresholds.The 50 µm front band approximated five cell diameters, while the 200 µm stroma cutoff approximated immune-infiltration attenuation.
  • 4.2 Feature extraction: Slides were classified as mass-forming, intermediate, or infiltrative using tumor-front fraction ϕ thresholds of ≤0.5, 0.5–0.8, and >0.8, respectively.A macro-tumor mask with disk radius ρ = 200 µm was used solely for the micro_interface_ratio quality-control metric.

4.3 Feature preprocessing

Feature preprocessing used log(1+x) transformation for 22 heavily right-skewed features and cohort-wide winsorization of all 38 features to reduce the influence of extreme values from segmentation artifacts.

  • Feature preprocessing: 22 heavily right-skewed features were log-transformed using log(1+x), including cell densities, ratios, distances, and heterogeneity measures.The heterogeneity measures included coefficients of variation.
  • Feature preprocessing: All 38 features were winsorized at the 0.5th and 99.5th percentiles computed per feature across the full cohort.This step mitigated extreme values arising from segmentation artifacts.

4.4 Survival analysis · 4.5 Molecular correlations

The study evaluated histomic associations with four survival endpoints using covariate-adjusted, cancer-stratified Cox models supplemented by assumption-free RMST analyses. It also assessed 38 histomic features against 293 molecular targets across cancer-specific and pan-cancer cohorts using unadjusted and adjusted rank-based correlations.

  • 4.4 Survival analysis: Four survival endpoints—OS, DSS, DFS, and PFS—were analyzed, with OS designated as the primary endpoint.Per-cancer-type results constituted the primary analyses, while pan-cancer models used cancer type as a stratification variable.
  • 4.4 Survival analysis: Covariate-adjusted survival models standardized continuous variables, one-hot encoded sex and stage, and used MICE with Bayesian ridge imputation.Pan-cancer stratified Cox regression allowed each cancer type its own baseline hazard while estimating a shared regression coefficient.
  • 4.4 Survival analysis: Proportional-hazards violations were assessed with Schoenfeld residuals, and associations with P < 0.05 were treated as non-proportional and invalidated for Cox inference.Invalidated P-values were set to 1.0 for BH correction, and RMST differences were computed as an assumption-free summary.
  • 4.4 Survival analysis: RMST used cancer-type-specific horizons: 1,095 days by default, 2 years for PAAD and MESO, and 5 years for THCA and PRAD.High- and low-feature groups were compared using 5,000 label permutations, with 95% bootstrap confidence intervals from 1,000 resamples.
  • 4.5 Molecular correlations: Correlations linked 38 histomic features to 293 molecular targets comprising gene expression, copy-number variation, Hallmark pathway scores, and immune cell scores.The 133 genes included curated cancer genes, immune checkpoints, and EMT/stemness markers.
  • 4.5 Molecular correlations: Unadjusted associations used Spearman correlations with analytical P-values and 95% bootstrap confidence intervals from 1,000 resamples.Adjusted associations used partial Spearman correlations after rank-transforming and residualizing features, targets, and covariates.
  • 4.5 Molecular correlations: The correlation analysis comprised 487,638 histomic–molecular pairs across 22 cohorts and two adjustment models, excluding combinations with insufficient data.Cancer types required n ≥30 samples, individual feature pairs required n ≥10 non-missing observations, and TSS was grouped into five frequent sites plus “Other”.

4.6 Categorical associations

Categorical associations between histomic features and molecular variables were tested separately by cancer type, using group-comparison methods for unadjusted analyses and rank-ANCOVA with permutation inference for covariate-adjusted analyses. Binary group ordering and minimum sample-size rules standardized effect interpretation and excluded sparse groups.

  • Categorical variables: Associations were tested separately for each cancer type across somatic mutations, copy-number alterations, and immune subtypes C1–C6.Mutation comparisons contrasted mutant versus wild-type, copy-number comparisons contrasted amplification/deletion versus neutral, and immune analyses covered subtypes C1–C6.
  • Unadjusted analyses: Unadjusted two-group comparisons used two-sided Mann–Whitney U tests with Cliff’s δ and 95% bootstrap confidence intervals from 1,000 resamples.Bootstrap intervals were used because tied values violate the continuity assumption underlying analytical Cliff’s δ intervals.
  • Covariate-adjusted analyses: Covariate-adjusted associations used rank-ANCOVA with Freedman–Lane permutation inference, comparing full and covariate-only models through an F-statistic.Responses and covariates were rank-transformed; P-values were obtained from 1,000 residual permutations.
  • Analysis conventions: Binary comparisons ordered mutant before wild-type and amplification/deletion before neutral, required n ≥30 total observations and ≥5 observations per group, and excluded groups smaller than 5.The deterministic ordering ensured consistent sign interpretation of Cliff’s δ.

4.7 Clustering

HistoAtlas clustered all 6,745 slides using 38-dimensional morphology features at pan-cancer and cancer-specific levels, selecting the pan-cancer solution through multiple validation criteria and assessing its stability by subsampling.

  • Clustering levels: 6,745 slides were clustered pan-cancer using the full 38-dimensional preprocessed feature vector, while cancer-specific clustering used within-cancer z-scores.Features underwent log transformation, winsorization, and z-scoring before L1 clustering.
  • Cluster selection: K = 10 was selected from candidates K ∈{3,...,25} because silhouette, Calinski–Harabasz, Davies–Bouldin, and gap statistic results converged at an interpretable granularity.Silhouette had a local maximum, Calinski–Harabasz a plateau, and Davies–Bouldin a local minimum at K = 10.
  • Stability assessment: 50 iterations of 80% subsampling without replacement assessed cluster stability using adjusted Rand index and mean best-match Jaccard index.Each iteration re-clustered with the same K and ninit = 10 and compared labels with the original solution.
  • Visualization: UMAP with nneighbors = 15, min_dist = 0.1, Euclidean metric, and random state 42 was used only to visualize the z-scored feature matrix.Clustering and statistical inference remained in the original 38-dimensional feature space.

4.8 Cluster enrichment · 4.9 Multiple testing correction

Cluster enrichment was evaluated with exact, rank-based, and permutation-based methods tailored to mutation, pathway, gene-set, and immune-subtype analyses. Multiple testing was controlled within biologically coherent families using Benjamini–Hochberg correction, with predefined survival-endpoint conventions and resampling safeguards.

  • 4.8 Cluster enrichment: Mutation enrichment used Fisher’s exact test on 2×2 mutated-versus-wild-type and in-cluster-versus-out-of-cluster tables, reporting odds ratios and exact 95% confidence intervals.The method was chosen for exact P-values without distributional assumptions, accepting its conservatism when only cluster membership is fixed.
  • 4.8 Cluster enrichment: Pathway activity distributions were compared between clusters with the Mann–Whitney U test, Cliff’s δ, and 1,000-resample bootstrap confidence intervals.A complementary GSEA analysis used Welch’s t-statistics comparing in-cluster with out-of-cluster expression.
  • 4.8 Cluster enrichment: GSEA significance used 1,000 phenotype-label permutations with separate positive and negative enrichment P-values, normalized enrichment scores, and pooled-null NES false discovery rates.Gene sets with FDR q < 0.25 were considered significant under the original GSEA convention.
  • 4.8 Cluster enrichment: Immune subtype enrichment for Thorsson C1–C6 classifications used Fisher’s exact test on subtype-presence and cluster-membership tables, with odds ratios, exact 95% confidence intervals, and observed-to-expected ratios.
  • 4.9 Multiple testing correction: All P-values underwent Benjamini–Hochberg correction within explicitly defined biologically coherent correction families rather than across unrelated analyses.Survival families were defined by cancer type, endpoint, and adjustment model.
  • 4.9 Multiple testing correction: Overall survival was the primary endpoint, while DSS, DFS, and PFS were treated as sensitivity analyses with independent correction by endpoint.The separation reflected partially overlapping event definitions that violate the independence assumption of joint correction.
  • 4.9 Multiple testing correction: Permutation P-values used the add-one correction P = (B+ + 1)/(B + 1), and bootstrap confidence intervals used 1,000 percentile resamples with reporting restricted to estimates having ≥90% valid resamples.Resamples were invalid when zero-variance columns or singular covariate matrices prevented model fitting.

4.10 Batch effect assessment · 4.11 Power analysis and evidence badges

HistoAtlas assessed tissue-source-site batch effects using variance decomposition, clustering diagnostics, and UMAP inspection, both globally and within cancer types. It also quantified minimum detectable effects at 80% power and assigned every association an evidence-strength badge based on statistical, effect-size, confidence-interval, and sample-size criteria.

  • 4.10 Batch effect assessment: PVCA decomposed top-principal-component variance into tissue source site, cancer type, and residual fractions using variance-weighted marginal ANOVA η2.The top components explained ≥80% cumulative variance, with up to 10 components.
  • 4.10 Batch effect assessment: PVCA was computed globally across all slides and separately within each cancer type to assess tissue-source-site variance.Feature matrices were standardized before PCA, and the three variance fractions were renormalized to sum to 1.0.
  • 4.10 Batch effect assessment: Silhouette analysis evaluated tissue source site as cluster assignments on standardized features using Euclidean distance and subsampling to 5,000 slides.Scores near zero or negative indicated minimal tissue-source-site clustering; scores above 0.25 and 0.5 flagged moderate and strong batch effects, respectively.
  • 4.10 Batch effect assessment: UMAP embeddings colored by tissue source site were visually inspected to confirm that slides did not cluster by source site after controlling for cancer type.Per-batch mean silhouette scores were also used to identify specific sites with elevated clustering.
  • 4.11 Power analysis and evidence badges: MDES was computed for every analysis at 80% power with α = 0.05 using analysis-specific methods for survival, correlations, and categorical associations.The methods were the Schoenfeld–Freedman approximation for survival, Fisher z-transform for correlations, and simulation-based power curves with 2,000 simulated datasets for categorical associations.
  • 4.11 Power analysis and evidence badges: Every atlas association received a badge—strong, moderate, suggestive, or insufficient—based on adjusted P values, effect-size thresholds, confidence intervals, and sample size.The sample-size criteria were n ≥100 for strong, n ≥50 for moderate, n ≥30 for suggestive, and n < 30 or missing statistics for insufficient evidence.

4.12 Web application · 4.13 Implementation and reproducibility

The HistoAtlas web atlas combines static, interactive pan-cancer visualizations with spatially traceable tissue and cell overlays, while its reproducible Python/Snakemake implementation, seeded analyses, and public code and data resources support reuse.

  • 4.12 Web application: Astro and React serve precomputed statistical results, feature profiles, cluster metadata, and visualization data as static JSON without runtime backend computation.The application supports pan-cancer and per-cancer views, including UMAP embeddings, feature distributions, and survival associations.
  • 4.12 Web application: Tissue compartment maps overlay five compartments across nine spatial zones, including tumor, stroma, necrosis, normal epithelium, and background.The interface provides spatial segmentation overlays per slide, exact whole-slide pixel coordinates, and a 10 µm scale bar.
  • 4.12 Web application: A toggleable overlay predicts 14 cell types in distinct colors and supports bidirectional navigation between statistical results, slides, tissue regions, features, and tiles.Users can trace survival hazard ratios and molecular correlations to feature pages and then to specific tiles on specific slides.
  • 4.13 Implementation and reproducibility: All analyses were implemented in Python 3.11 and orchestrated through a Snakemake directed acyclic graph of computational dependencies.The workflow defines the computational dependencies used across the analysis pipeline.
  • 4.13 Implementation and reproducibility: Explicit random seeds were used for permutation tests, bootstrap confidence intervals, cluster-stability subsampling, and K-means initialization.The reported library versions include lifelines 0.29.0, scipy 1.12.0, scikit-learn 1.4.0, and umap-learn 0.5.5.
  • 4.13 Implementation and reproducibility: Analysis code, precomputed results, and the web application are publicly available, with the atlas accessible at histoatlas.com and 6,745 slide UUIDs listed in the repository.Derived feature matrices, precomputed statistical results, and model weights are planned for deposition in a public repository with a persistent DOI before publication.

Histomic category

HistoAtlas organizes 38 downstream histomic features into composition, density, morphology, spatial, heterogeneity, and ratio categories, linking morphology to molecular programs and distinct immune and clinical archetypes. The atlas also identifies tissue morphological heterogeneity as protective in adjusted analyses across most evaluated cancers.

  • Molecular programs: Morphological features recapitulated molecular programs across 21 cancer types through correlations with 50 Hallmark pathway scores.The analysis used mean Spearman correlations across cancer types in an unadjusted model.
  • Morphological archetypes: Morphological clusters mapped to immune subtypes and distinct molecular archetypes, including immune-hot clusters enriched for C2 IFN-γ dominance.Clusters 3, 4, and 7 were all C2-positive across pan-cancer analyses.
  • Feature categories: 38 histomic features were organized into composition, density, morphology, spatial, heterogeneity, and ratio categories for downstream analyses.Supplementary Table 1 defines 40 extracted features, of which 38 were used downstream.
  • Clinical associations: Tissue morphological heterogeneity was protective in 11/15 cancers after adjustment, contradicting the tissue heterogeneity paradigm.The biological plausibility audit classified this claim as contradicted, reflecting a category distinction.
  • Clinical associations: Two distinct immune-cold phenotypes had opposite survival outcomes, while proliferative clusters showed worse survival with r ≈0.85 across clusters.These cluster-level findings were reported as well-established in the biological plausibility audit.

Supplementary Note 2: Cross-Endpoint Replication of Survival Associations

Cross-endpoint analysis tested whether 60 OS-significant feature–cancer associations replicated for DSS, PFS, and DFS with matching hazard-ratio direction and FDR < 0.05. DSS replication was higher than PFS but is inflated by overlapping event definitions, whereas direction concordance remained complete among significant replications.

  • Replication framework: 60 OS-significant feature–cancer pairs were evaluated for replication in DSS, PFS, and DFS using matching hazard-ratio direction and FDR < 0.05.DFS was excluded from per-pair counts because it was unavailable in the concordance output.
  • Replication rates and limitations: 62.7% DSS replication exceeded 53.3% PFS replication, but DSS concordance overstates biological cross-validation because the endpoints share overlapping event definitions and follow-up windows.DSS differs from OS mainly through censoring of non-cancer deaths, making the two endpoints non-independent.
  • Direction concordance: All DSS- and PFS-replicated associations agreed with OS in hazard-ratio direction.When significance thresholds were relaxed, a larger fraction pointed in the same direction across endpoints, consistent with biological signal attenuated by fewer secondary-endpoint events.
  • Feature-level patterns: Strongest OS effects showed the highest cross-endpoint replication, including intratumoral lymphocyte density in BRCA, HNSC, and pan-cancer analyses.Intratumoral apoptotic index and tumor pleomorphism index replicated across all three endpoints in LIHC, while mitotic index replicated for DSS and PFS in MESO.
Loading 2603.16587v1…