Source-linked AI summary
CHIMERA Challenge Task 2 and 3: Response Subtypes Classification and Progression Survival Prediction in Bladder Cancer Patients using Multimodal Datasets
Catherine Chia, Tongjie Wang, Robert Spaans, Maryam Mohammadlou, Farbod Khoraminia, J. Alberto Nakauma-González, Adam Kowalewski, Parandzem Khachatryan, Domingos Oliveira, Khrystyna Faryna, CHIMERA Challenge Consortium, Marlies Wakkee, Sita Vermeulen, Tahlita Zuiverloon, Nadieh Khalili
TL;DR
HR-NMIBC outcomes remain heterogeneous, while existing clinical stratification and molecular subtyping have limited individual or routine clinical availability. CHIMERA benchmarked multimodal prediction across response-subtype and progression tasks, with hidden-test performance reaching a weighted F1 of 0.7261 and a C-index of 0.6828.
Problem
HR-NMIBC outcomes are heterogeneous, while traditional clinicopathological risk groups and RNA-seq subtyping have limited individual and routine clinical availability.
Method
CHIMERA established a standardized multimodal benchmark combining H&E whole-slide images, structured clinicopathological data, and RNA sequencing with public training and hidden evaluation.
Results
Performance varied across tasks and cohorts; hidden-test results reached a weighted F1 of 0.7261 for Task BRS.
Takeaways & Limitations
The benchmark identifies cohort variation and missing structured data as practical considerations for multimodal prediction.
Takeaways & Limitations
Progression analyses were constrained by few events, making cohort-specific C-indices, ablation effects, and patient-level error patterns uncertain and potentially unstable.
Abstract
from arXiv · showhide
High-risk non-muscle-invasive bladder cancer (HR-NMIBC) carries substantial risks of recurrence and progression, while current clinical risk stratification remains limited. CHIMERA was established as a multimodal AI challenge to benchmark prediction in HR-NMIBC under standardized evaluation. Task BRS predicts RNA-seq-defined BCG Response Subtypes from histopathology and structured clinicopathological data, whereas Task Progression models time-to-progression using histopathology, structured data, and RNA sequencing. A multimodal dataset of 368 patients was divided into public training and hidden validation and test sets. In total, 159 submissions were made, and 13 top-performing models were selected for benchmarking. The best models achieved a weighted F1 score of 0.73 for Task BRS and a C-index of 0.68 for Task Progression. Post-challenge analyses revealed task-dependent modality contributions, cohort-dependent performance degradation, and sensitivity to missing structured data. In Task BRS, histopathology partly compensated for pathology-derived structured variables, whereas progression models showed greater dependence on complementary inputs. Cross-model error analysis further identified patients that were consistently difficult across different architectures, with T1 substage associated with prediction difficulty. These findings highlight barriers to transportability and the importance of missingness-aware modeling and independent multi-institutional validation. CHIMERA provides a standardized multimodal benchmark for bladder cancer and a framework for studying not only model performance, but also robustness, information sufficiency, and patient-level prediction failure.
Highlights
CHIMERA establishes a standardized multimodal bladder cancer benchmark, with top models achieving weighted F1 of 0.73 for Task BRS and a C-index of 0.68 for Task Progression. Findings also show limits from cohort shift and missing data, while cross-model errors identify T1-associated consensus-hard patients.
- Highlights: CHIMERA provides the first standardized multimodal benchmark for bladder cancer.
- Highlights: Cohort shift and missing data limit multimodal model generalization.
- Highlights: Cross-model errors reveal consensus-hard patients associated with T1 substage.
- Highlights: 0.73 weighted F1 for Task BRS and 0.68 C-index for Task Progression were achieved by 13 top-performing models selected from 159 submissions.Benchmarking also revealed task-dependent modality contributions: histopathology compensated for missing pathology-derived clinicopathological variables in Task BRS, whereas Task Progression models depended on such variables.
1. Introduction
HR-NMIBC has substantial recurrence and progression after BCG, while clinically relevant molecular subtyping remains difficult to deploy routinely. CHIMERA addresses this translational and benchmarking gap through standardized multimodal prediction tasks spanning BCG response biology and progression.
- Introduction: ˜50% of HR-NMIBC patients experience recurrence and/or progression after TURBT followed by BCG, driving intensive surveillance and difficult treatment-escalation decisions.
- Introduction: Molecularly defined BCG Response Subtypes identify biologically distinct outcomes, but RNA-seq subtyping remains time- and cost-intensive and is not broadly embedded in routine pathology workflows.BRS3 was associated with inferior recurrence-free and progression-free survival and BCG failure.
- Introduction: Robust multimodal assessment remains difficult because comparisons face heterogeneous cohorts, endpoints, modalities, evaluation protocols, incomplete inputs, and institution-specific acquisition effects.
- Introduction: CHIMERA was established as an international multimodal benchmark for predicting RNA-seq-based BCG response subtypes and patient-level progression from histopathology, structured data, and molecular inputs.Task BRS uses H&E whole-slide images and structured patient and clinicopathological data, whereas Task Progression evaluates time-to-progression.
- Introduction: CHIMERA provides a shared benchmark spanning morphology, molecular data, and clinically relevant outcomes to support research on personalized NMIBC management.The authors describe it as the first challenge in digital pathology and multimodal analysis for bladder cancer.
2. Material and methods
CHIMERA used multimodal HR-NMIBC tumor cases combining structured clinicopathological data, RNA-seq, and digitized H&E slides. Task BRS used RNA-seq-derived BRS as the reference standard, whereas Task Progression predicted patient-level progression risk using progression status and follow-up time evaluated by censored concordance index.
- Cohort and data construction: The dataset defined each case through structured data, RNA-seq, and digitized H&E slides, excluding cases missing a modality and Urolife cases classified as intermediate risk.The remaining multimodal cases were split into training, validation, and test sets.
- Data harmonization: Structured variables were harmonized across cohorts, while unavailable BCG instillation counts and other missing values were encoded as -1.Urolife stage and T1 substage labels were converted to corresponding harmonized categories before modeling.
- Endpoint definition: Progression required advancement to muscle-invasive disease, regional lymph-node metastasis, or distant metastasis, excluding treatment failure, discontinuation, or cystectomy without documented muscle-invasive disease.The progression endpoint was therefore related to, but not equivalent to, BCG treatment failure.
- Task definitions and evaluation: For Task BRS, RNA-seq-derived BRS served as the reference standard; for Task Progression, BRS was an input and progression status plus time to progression or follow-up end formed the reference standard.
- Data processing: Histopathology slides were digitized at 0.25 µm/pixel, while RNA-seq counts underwent DESeq2 transformation, normalization, and variance stabilization with only protein-coding genes retained.
- Task definitions and evaluation: Censored concordance index measured Task Progression performance by the proportion of comparable patient pairs correctly ordered by predicted risk.Predicted risk was defined as the negation of predicted time to progression, enabling comparison of survival models by patient risk ranking.
3. Results
Across hidden test evaluation, performance ranged from 0.60 to 0.73 for Task BRS and from 0.57 to 0.68 for Task Progression. Results also showed cohort- and modality-dependent robustness, with structured-data masking affecting progression more strongly and recurring patient-level prediction errors.
- Challenge performance: Hidden test estimates ranged from 0.60 to 0.73 for Task BRS and from 0.57 to 0.68 for Task Progression.Team-reported development estimates were higher, ranging from approximately 0.63 to 0.85 and 0.73 to 0.91, respectively.
- Random masking: Progression models deteriorated with increasing random masking in Erasmus Cohort B, whereas Urolife showed lower baselines and less monotonic responses.For example, HKKH fell from 0.757 to 0.486 at 50% masking in Erasmus Cohort B, while Urolife HKKH remained near baseline at 25% and 50%.
- Structured-data sensitivity: Structured-variable masking generally had small effects on Task BRS but larger effects on Task Progression, including NMIL falling from 0.699 to 0.410 after masking BRS in Cohort B.In Urolife, NMIL also decreased from 0.479 to 0.319 after masking sex and to 0.326 after masking substage.
- Modality ablation: Removing image-derived variables produced small or inconsistent changes, with no corrected comparison significant for either task.For Task Progression, full models were numerically higher by 0.031 to 0.092 in some settings, but all 95% confidence intervals crossed zero and all adjusted p-values were at least 0.670.
- Patient-level errors: Cross-model analysis identified recurring patient-level failures, including five patients misclassified by all eight BRS models and T1 substage differences between hard and easy patients.The analysis identified 41 hard and 35 easy BRS patients; progression analysis identified 24 hard patients using failures across at least three of five models.
4. Discussion
CHIMERA established a standardized multimodal benchmark for BCG response-subtype and progression prediction in HR-NMIBC, achieving weighted F1 0.7261 and C-index 0.6828 on hidden tests. Discussion analyses highlighted cohort-dependent transportability, task-specific modality contributions, missing-data challenges, and the need for cautious interpretation and external validation.
- Overall benchmark: 0.7261 weighted F1 for Task BRS and 0.6828 C-index for Task Progression were the highest hidden-test performances.These results summarize CHIMERA’s benchmark performance across its two clinically relevant prediction tasks.
- Cohort transportability: Progression models performed better in Erasmus Cohort B than Urolife, demonstrating cohort-dependent transportability despite competitive performance by some methods.The analyses identified cohort shift across structured variables, H&E features, and transcriptomic representations, but did not establish which modality drove performance loss.
- Cohort transportability: Adding RNA-seq to H&E did not significantly change discrimination in Erasmus Cohort B and only nominally improved Urolife performance after correction.Because this comparison involved one voluntarily supplied post-hoc model pair, it was hypothesis-generating rather than a challenge-wide estimate of RNA-seq contribution.
- Modality contributions: Task BRS showed small average changes after masking individual structured variables, whereas Task Progression was more sensitive to clinical, treatment-course, and BCG-response information.Progression models also showed sensitivity to BRS in some architectures, suggesting potential value from predefined biological representations, although this was not uniform.
- Pathology-derived variables: Pathology-derived structured variables showed limited measurable incremental endpoint value beyond H&E in evaluated models, particularly for Task BRS, but this does not justify eliminating structured curation.The reduced-model and ablation findings were limited by statistical power, and correlated inputs or fusion weighting could explain the observed redundancy.
- Missing data and interpretation: Substantial structured-data missingness, including complete absence of no_instillations in Urolife, made missing-data management a practical benchmark challenge.Most participants used mean or median imputation, while only the WL Team implemented learned imputation; ablation results should therefore be interpreted as sensitivity to both information removal and missingness patterns.
5. Conclusions
CHIMERA established a standardized, reproducible multimodal benchmark for HR-NMIBC, achieving strong benchmark performance while revealing architecture- and cohort-dependent modality value. The framework supports evaluation of missing-input robustness and patient-level failure modes but does not establish clinical readiness.
- Conclusions: 0.726 weighted F1 for BRS prediction and 0.683 C-index for progression prediction were the best test-set performances in CHIMERA’s reproducible multimodal benchmark.The benchmark combined H&E whole-slide images, structured clinicopathological information, and RNA sequencing within the same HR-NMIBC population.
- Conclusions: Modality contributions depended on architecture and cohort, while BRS models generally tolerated removal or partial masking of individual structured variables.Reduced BRS models retained performance after image-derived variables were removed, supporting partial informational overlap without demonstrating direct recovery of pathology variables.
- Conclusions: CHIMERA does not establish clinical readiness or universal benefits from additional modalities, but provides a reproducible framework for studying cohort shift, missingness robustness, calibration, and failure modes.Future work should extend the framework to larger multi-institutional cohorts, controlled modality baselines, missingness-aware modeling, and domain alignment.
CRediT authorship contribution statement
The CRediT statement assigns authorship contributions across conceptualization, data curation, analysis, methodology, software, supervision, project administration, and writing. It identifies varied combinations of these roles for the named contributors.
- Catherine Chia contributed across nearly all listed roles, including conceptualization, data curation, formal analysis, methodology, software, validation, visualization, project administration, and writing.
- Tongjie Wang contributed to conceptualization, formal analysis, investigation, methodology, software, validation, visualization, and both original-draft and review writing.
- Robert N. Spaans, Tahlita Zuiverloon, and Nadieh Khalili are credited with combinations of conceptualization, resources, supervision, software, project administration, funding acquisition, and writing.
Ethics Approval
The study received ethics approval from the relevant Dutch medical and human research committees, and treatment and follow-up followed specified clinical protocols and guidelines.
- Ethics Approval: Ethics approval covered the Erasmus MC and Arnhem-Nijmegen committees, while BCG treatment followed SWOG and clinical follow-up followed EAU guidelines.The UroLife and Nijmegen Bladder Cancer Study cohorts were approved by the Arnhem-Nijmegen Committee for Human Research.
Funding
The research was funded by Hanarth Fond, the Dutch Research Council/NWO Vidi project AIPRECISE, and the European Union’s Horizon Europe CLARIFY project; Astellas Pharma sponsored the CHIMERA challenge prizes.
- Funding: Funding came from Hanarth Fond, NWO Vidi project AIPRECISE, and the EU Horizon Europe CLARIFY project, while Astellas Pharma sponsored the CHIMERA challenge prizes.The grants supported named researchers, including Marlies Wakkee, Catherine Chia, Tongjie Wang, Tahlita Zuiverloon, and Farbod Khoraminia.
Figure captions
The figures depict the CHIMERA challenge workflow, multimodal task data distribution, and patient-level error analysis for progression prediction. Patient-level difficulty was associated with T1 substage in the progression analysis.
- Challenge workflow: The figures summarize CHIMERA’s workflow from public training through hidden validation and test evaluation, followed by benchmarking of top-performing solutions.The challenge covered Task BRS and Task Progression in bladder cancer multimodal modeling.
- Patient-level error analysis: Patient-level progression errors were visualized across models, with patients grouped by prediction difficulty and T1 substage remaining significant after correction (p < 0.001).The heatmap marks model-level failures and patients failed by all models, alongside outcome and histology tracks.
Supplementary Material … Task BRS - Team BioToTem - Multimodal model wih ACMIL
The BioToTem team developed a multimodal BRS framework combining histopathology with clinical data. Its preprocessing used augmented H&E patches, standardized clinical variables, and an explicit category for missing values.
- S1. CHIMERA Challenge Organizers: Task 2 (BRS) and 3 (Progression): The challenge organizer group included researchers from pathology, urology, dermatology, and related departments across Dutch medical institutions.
- Task BRS - Team BioToTem - Multimodal model wih ACMIL: BioToTem integrated histopathology and clinical data in a multimodal framework for Task BRS.
- Task BRS - Team BioToTem - Multimodal model wih ACMIL: H&E whole-slide images were analyzed at 0.5 mpp using provided tissue masks and 224 x 224 patches.
- Task BRS - Team BioToTem - Multimodal model wih ACMIL: Data augmentation was used to improve generalization instead of stain normalization.
- Task BRS - Team BioToTem - Multimodal model wih ACMIL: Clinical categorical variables were one-hot encoded, while continuous variables were normalized to zero mean and unit variance.
- Task BRS - Team BioToTem - Multimodal model wih ACMIL: Missing clinical values were retained as a dedicated unknown category to capture potential missingness-related signals.
S4.1. Task BRS - Team GRIS - Multimodal model with multiclass predictions … S4.8. Task Progression - Team SMILE - Unimodal using a discrete-time survival
The selected teams used diverse multimodal and unimodal strategies for BRS classification and progression survival prediction, varying histopathology encoders, clinical-data representations, fusion schemes, and missing-data handling. The progression approach additionally evaluated clinical and RNA-based survival configurations with split-specific preprocessing.
- S4.1. Task BRS - Team GRIS - Multimodal model with multiclass predictions: GRIS combined H&E whole-slide features and 23 missingness-aware clinical features for multiclass BRS prediction.Slides were tissue-segmented, processed at multiple resolutions with CLAM and Macenko normalization, and encoded using UNI2-h.
- S4.2. Task BRS - Team BUAA_REMEX - Multimodal model using PANTHER: BUAA_REMEX integrated histopathology with clinical data converted into text, representing missing values as “not available.”Histopathology used CONCHV1_5 and PANTHER features, while clinical text used the CONCH text encoder and a modified PANTHER.
- S4.2. Task BRS - Team BUAA_REMEX - Multimodal model using PANTHER: BUAA_REMEX also applied slide2vec preprocessing, augmentation, median imputation, and UNI-based 256 x 256 patch embeddings for its multimodal BRS model.Gaussian noise and patch dropout were used for whole-slide images, while Gaussian noise was applied to selected clinical variables.
- S4.4. Task BRS - Team WL - Unimodal model using AutoGluon-Tabular: WL trained unimodal clinical-data BRS models with AutoGluon-Tabular and 10-fold stratified cross-validation.The base learners included XGBoost, LightGBM, CatBoost, RandomForests, RealMLP, and TabM, with listwise deletion for samples missing reference standards.
- S4.6. Task BRS - Team HKKH - multimodal model using MADMIL: HKKH fused UNI-encoded H&E slide embeddings with 30-dimensional clinical embeddings before BRS classification.Slides were processed at 1.0 mpp with 224 x 224 patches, and slide-level representations were generated using multi-head attention-based deep multiple instance learning.
- S4.6. Task BRS - Team HKKH - multimodal model using MADMIL: Aillis used a unimodal histopathology strategy with density-based selection of dark patches and either EfficientNet-B5 or UNI2-H encoding.The approach used 0.5 mpp H&E slides, multiple patch-grid configurations, Albumentations augmentation, and excluded four problematic slides.
- S4.8. Task Progression - Team SMILE - Unimodal using a discrete-time survival: SMILE developed progression survival models using clinical data and RNA in unimodal and multimodal configurations.Clinical variables were one-hot encoded and normalized, missing values were imputed by type, and preprocessing was fitted on training data before application to validation and test sets.
S4.9. Task Progression - Team HKKH - Multimodal by evaluating three fusion … S6. Additional Robustness analysis
The progression analyses evaluated multimodal and clinical-only survival models, while robustness analyses examined reduced structured-data effects in Task BRS and cohort-specific modality contributions in Task Progression. TIA-Pegasus’s multimodal model achieved an external C-index of 0.574 despite higher internal performance.
- S4.9. Task Progression - Team HKKH - Multimodal by evaluating three fusion: HKKH developed multimodal survival models that fused clinical data, histopathology, and RNA-seq through multiple fusion strategies.Whole-slide images were processed at 1.0 mpp with UNI features and gated-attention ABMIL aggregation, while RNA data used a two-layer feedforward encoder.
- S4.11. Task Progression - Team WL - Unimodal using ensemble of models: The WL team built a clinical data-only survival model with imputation, standardization, one-hot encoding, and 10-fold stratified cross-validation.The approach excluded patients without survival outcomes and used rare-bin guarding during stratification.
- S4.12. Task Progression - Team TIA-Pegasus - Multimodal using ModalSurv: TIA-Pegasus compared unimodal and multimodal survival frameworks using histopathology, clinical variables, and RNA expression reduced to 128 dimensions.Histopathology was encoded with CONCH and aggregated with Titan into 768-dimensional slide embeddings.
- S4.12. Task Progression - Team TIA-Pegasus - Multimodal using ModalSurv: 0.574 external C-index was achieved by TIA-Pegasus’s submitted multimodal survival model, versus 0.733 internally.The clinical-only unimodal model had a higher internal C-index of 0.749, but the multimodal model was submitted.
- S5. Extended Task BRS analysis: Task BRS robustness analysis compared full and reduced structured-data models across the combined test set and Erasmus Cohort B and Urolife Cohort U.Table S1 reports AUC, ΔAUC, adjusted p-values, and 95% percentile bootstrap confidence intervals for the three teams submitting reduced models.
- S6. Additional Robustness analysis: Additional Task Progression robustness analysis compared TIA-Pegasus H&E-only and H&E+RNA-seq sub-models by cohort using C-index.Figure S2 reports 95% percentile bootstrap confidence intervals, adjusted within-cohort p-values, and a chance-level C-index of 0.5.