Source-linked AI summary

The Ceiling Is in the Channel: Auditing Learner Gaps and Measurement Frontiers in Clinical Prediction

Sayeed Shafayet Chowdhury, Nusrat Jahan, Snehasis Mukhopadhyay, Shiaofen Fang, Vijay R. Ramakrishnan

arXiv:2609.01909v1cs.AI

TL;DR

Clinical prediction may saturate because learners fail to extract recorded information or because the measurement channel imposes a population frontier. The paper formalizes and audits this distinction, finding recurring learner–channel patterns across cohorts and a broader synthesis of clinical tasks, with measurement changes often extending performance when learner headroom is small.

  • Problem

    Clinical prediction studies often conflate limited learner extraction with a limiting recorded measurement channel, making the source of saturation difficult to distinguish.

  • Method

    The paper characterizes balanced-accuracy frontiers through total variation and estimates them with cross-fitted posteriors, permutation-null and underfit diagnostics, cohort audits, and a PRISMA-guided synthesis.

  • Results

    Across three cohorts and 104 additional clinical tasks, well-tuned boosting nearly reaches estimated frontiers while same-channel gains diminish and richer or complementary measurement channels often extend performance.

  • Takeaways & Limitations

    A large learner gap supports model improvement, whereas a small gap shifts attention toward measurement, labels, and deployment-relevant channel changes.

  • Takeaways & Limitations

    The evidence synthesis is descriptive with heterogeneous metrics and validation designs, and the observational cohorts do not establish that changing a measurement causally improves outcomes.

Abstract

from arXiv · show

Clinical prediction can saturate for two different reasons: a fitted learner may fail to extract available information, or the recorded variables may impose a population frontier. We separate these quantities through the \emph{learner gap} and the \emph{measurement-channel ceiling}. Optimal balanced accuracy is characterized by total-variation separation, yielding architecture invariance, a sharp partial-identification result under replacement contamination, a cross-fitted ceiling estimator, and exact conditions for multimodal decision improvement. We add two finite-sample diagnostics, namely a label-permutation optimism floor and an underfit curve, and validate the audit on three real cohorts: UCI readmission ($n=99{,}343$), BRFSS diabetes ($n=253{,}680$), and NHANES HbA1c ($n=10{,}219$). Well-tuned gradient boosting nearly reaches the estimated frontier in UCI and BRFSS, whereas deliberately or practically deficient learners retain large gaps. NHANES yields a null difference between questionnaire and measured marginal frontiers but a significant joint complementarity gain, refining the simplistic claim that an objective modality must dominate. Across all cohorts, modest AUROC gains coexist with substantially larger Bayes decision-flip rates, and several architectures estimate similar frontiers while their achieved balanced accuracy differs sharply. A PRISMA-guided synthesis of 104 clinical tasks then shows that the same channel-level regularities recur across more than 18 disease categories: a broad but non-universal structured-clinical region, diminishing same-channel gains across model families, and higher performance when measurement channels change. The framework converts saturation from an empirical observation into an auditable decision: improve the learner when headroom remains; improve measurement when it does not.

Introduction

The paper distinguishes learner limitations from fixed measurement-channel limits in clinical prediction. It formalizes the distinction through balanced-accuracy frontiers, cross-fitted auditing, and diagnostics for estimator bias and multimodal improvement.

  • Structured-record models often report AUROC near 0.78–0.88 despite increasing capacity and cohort size.
  • The measurement-channel ceiling is the population Bayes frontier from observed variables, while the learner gap is the difference between that frontier and achieved performance.More data, optimization, or architectural richness may close the learner gap but cannot raise a fixed-channel frontier.
  • The framework estimates fixed-channel frontiers with out-of-fold equal-prior posteriors and audits them using permutation-null and underfit diagnostics.Random labels should yield a ceiling of 0.5, while a positive final underfit-curve increment indicates a lower-bound estimate.
  • Optimal balanced accuracy equals total-variation separation of class-conditional distributions, making the population frontier architecture-invariant under data processing.Finite learners can still differ in how closely they approach this common upper bound.
  • Additional modalities strictly improve hard classification exactly when Bayes decisions change on a positive-probability set; positive conditional mutual information alone is insufficient.The framework therefore distinguishes confidence refinement from actual decision improvement.
  • Under shared replacement contamination, a ceiling of 0.85 identifies separation 0.70 but permits only α ∈[0, 0.30].The upper endpoint is a limiting compatible contamination level, not an estimated clinical noise rate.

Controlled Validation

Controlled experiments test the audit against analytically known frontiers across contamination, multimodal complementarity, repeated measurements, and cross-fitted recovery. Flexible nonlinear learners approach the frontier, while logistic regression can remain below it when its decision class is misspecified.

  • Replacement contamination: At α = 0.30, the replacement-contamination experiment has an exact controlled ceiling of 0.85.The envelope is CBA = 1 −α/2 when latent classes have disjoint nonlinear supports and replacement is class-independent.
  • Gaussian complementarity: The Gaussian joint frontier is 0.9374, exceeding marginal ceilings of 0.80 for X and 0.90 for Z.This demonstrates complementary information beyond the stronger marginal channel.
  • Reliability and repetition: Repeated measurements raise the frontier along saturating curves for ρ ∈{0.30, 0.60, 0.90} but cannot exceed the latent-score ceiling 0.95.This panel is a prospective theoretical prediction rather than an empirical validation in the three real cohorts.
  • Learner validation: Flexible nonlinear learners approach the known frontier, whereas logistic regression remains below it because its decision class cannot express the radial boundary.
  • Cross-fitted recovery: The cross-fitted estimator closely recovers the known contamination frontier across α ∈{0, .1, . . . , .5}.Panels (a), (b), and (d) validate the audit mechanisms; panel (c) remains prospective.

Real-Cohort Frontier Audits

The cohort audits estimate fixed-channel balanced-accuracy frontiers, learner gaps, and complementarity across UCI, BRFSS, and NHANES. Well-tuned boosting nearly reaches the UCI and BRFSS frontiers, while channel combinations and learner choice produce distinct decision outcomes.

  • UCI readmission: UCI’s ceiling is 0.6225 and best balanced accuracy is 0.6223, yielding G = +0.0002.The underfit sequence ends at 0.6225 with final increment +0.0029, so the ceiling is reported as a lower bound.
  • BRFSS: BRFSS’s ceiling is 0.7522, best balanced accuracy is 0.7518, and G = +0.0003.Its underfit sequence is treated as converged, with final increment −0.0002.
  • NHANES: NHANES questionnaire and measured frontiers differ by +0.0042 with overlapping intervals, whereas the joint frontier reaches 0.7623 and complementarity is +0.0471.The joint interval is disjoint from the measured-channel interval, supporting complementary rather than modality-dominant information.
  • Ranking and decisions: Across channels, AUROC gains of +0.0617, +0.0672, and +0.0705 accompany decision-flip rates of 0.2996, 0.1907, and 0.2105.The reported decision changes are 4.9×, 2.8×, and 3.0× larger than the corresponding ranking gains.
  • Architecture comparison: Across BRFSS and NHANES, estimated frontiers span 0.0087 and 0.0195 across architectures, while achieved balanced accuracy spans 0.173 and 0.182.This shows that similar fixed-channel frontiers can coexist with sharply different learner performance.

Large-Scale Empirical Observations Across

The broader synthesis finds recurring but non-universal channel-level patterns across 104 clinical tasks and more than 18 disease categories. Its descriptive evidence links diminishing same-channel model gains with higher performance after channel expansion, while emphasizing heterogeneous designs and limits on causal interpretation.

  • Synthesis scope: The PRISMA-guided synthesis covers 104 task-level observations spanning more than 18 disease categories without pooling them as a common estimand.Outcomes, horizons, prevalence, validation, and metrics differ across observations.
  • Cross-domain recurrence: AUROC values repeatedly intersect a region near 0.78–0.88 across surgical, cardiovascular, obstetric, endocrine/renal, neurological, and oncological tasks.The synthesis also reports counterexamples, including lower readmission and chronic-pain tasks and higher ECG, imaging, and Parkinson prediction tasks.
  • Learner saturation: Most model-family gains occur before strong nonlinear tabular learners, while imaging CNNs operate on a different measurement channel.Within cohorts, frontier estimates span only 0.0087 and 0.0195 across architectures, versus achieved balanced-accuracy spans of 0.173 and 0.182.
  • Channel expansion: Six comparisons report higher performance after multimodal channel expansion, with clinical-only values near 0.75–0.83 and multimodal values near 0.83–0.90.These are descriptive heterogeneous summaries rather than pooled frontier estimates.
  • Limitations: The synthesis is descriptive because reviews overlap, metrics and validation designs differ, and patient-level uncertainty is often unavailable.The observational cohorts do not prove that changing a measurement causally improves outcomes.
  • Reporting implications: A benchmark should report achieved balanced accuracy, the cross-fitted frontier, gap G, the permutation-null floor, and the underfit verdict.The gap measures extractive headroom, while the diagnostics assess optimism and frontier stability.
  • Reporting implications: Clinical studies should separate ranking, decisions, and calibration by reporting AUROC, a prevalence-robust decision metric, calibration, and threshold-selection protocol.The paper notes that added channels may move probabilities across decision boundaries even when AUROC gains are modest.
  • Experimental design: When headroom is small, the next experiment should test measurement change through repeated administration, adjudicated outcomes, temporal features, or complementary modalities.A frontier shift should be demonstrated with uncertainty on the frontier difference and a decision-flip analysis.

Conclusion

Clinical prediction has two scaling problems: learners determine how closely models approach recorded information, while measurement channels determine population frontiers. Total-variation theory and cross-fitted audits make this distinction measurable across cohorts and channel settings.

  • Conclusion: The learner determines extraction from recorded information, whereas the measurement channel determines the population frontier.The framework therefore distinguishes learner gaps from measurement-channel ceilings.
  • Conclusion: Cross-fitted audits with permutation and underfit diagnostics quantify whether saturation reflects learner limitations or the recorded channel.Across three cohorts, well-tuned boosting has near-zero gaps while deficient learners retain large gaps.
  • Conclusion: Channel complementarity can change clinical decisions even when ranking gains are modest.The paper frames this distinction as a practical basis for deciding between model improvement and measurement improvement.

Background

The supplement provides proofs, audit and cohort details, the PRISMA-guided synthesis protocol, and additional empirical figures. These materials support the paper’s theory, experiments, and source meta-analysis.

  • Supplement contents: The supplement contains complete proofs for the main paper’s lemmas, theorems, and propositions.It also includes methodological details for the cross-fitted frontier audit and real-cohort experiments.
  • Supplement contents: It provides the complete PRISMA-guided evidence-synthesis protocol, descriptive tables, and additional empirical figures.Each additional figure is accompanied by detailed interpretation and methodological qualification.

A Channel-Ceiling Theory of Clinical Prediction

The paper separates information imposed by the observed measurement channel from information left unextracted by finite learners. It characterizes balanced-accuracy ceilings through total-variation separation, then extends the framework to contamination, cross-fitted estimation, repeated measurements, and multimodal prediction.

  • Foundational distinction: Clinical performance reflects both class information in observed variables and a learner’s ability to extract it.The channel determines a population frontier, while finite-sample estimation, optimization, and model class determine the learner gap.
  • Foundational distinction: Optimal balanced accuracy is characterized by total-variation separation of the class-conditional distributions.The optimum is attained by the equal-prior likelihood-ratio rule.
  • Architecture invariance: All representations computed from the same observed variables share a common population upper bound, although finite learners can approach it by different amounts.Data processing yields architecture invariance at the frontier, not equal achieved performance across models.
  • Replacement contamination: A ceiling of 0.85 identifies effective separation 0.70 but permits replacement contamination α ∈[0, 0.30].Under shared replacement contamination, the ceiling depends on observed separation through κX = (1−α)τ, so the contamination level is only partially identified.
  • Operational audit: The cross-fitted ceiling estimator uses out-of-fold equal-prior posterior predictions to estimate the population frontier without evaluating posterior models on training observations.The paper recommends multiple flexible posterior learners, nested cross-validation, and bootstrap intervals because finite-sample underfitting can bias the estimate toward 1/2.
  • Measurement improvement: Repeated measurements are predicted to raise the ceiling along a reliability-determined saturating curve, while an added modality strictly improves raw accuracy only when it changes Bayes decisions on a positive-probability set.Positive conditional mutual information alone may refine posterior confidence without crossing the decision threshold; complementary modalities can nevertheless raise the joint balanced-accuracy ceiling above both marginal ceilings.

Cross-Fitted Frontier Audit

The audit estimates a fixed-channel frontier from equal-prior, out-of-fold posteriors and separates it from achieved balanced accuracy. Permutation and training-fraction diagnostics assess optimism and convergence before interpreting the estimate.

  • Posterior Targeting: The equal-prior posterior is obtained from an ordinary prevalence-weighted posterior by posterior transformation, or directly through balanced class weighting.The transformation targets the equal-prior decision problem rather than the prevalence-weighted marginal distribution.
  • Cross-Fitted Audit: Cross-fitting fits a probabilistic learner on other folds, stores out-of-fold equal-prior posteriors, and computes the frontier, achieved balanced accuracy, and learner gap.The protocol respects patient groups when repeated observations belong to one patient and bootstraps the independent sampling unit.
  • Bias Diagnostics: The permutation-null diagnostic treats the true ceiling as 0.5 after valid label shuffling and uses the excess as a finite-sample optimism floor.A positive value measures posterior plug-in overconfidence, but subtracting it is not guaranteed to remove bias under the original signal distribution.
  • Convergence Diagnostics: The underfit curve refits the learner at training fractions 0.25, 0.5, 0.75, and 1.0 while preserving the evaluation protocol.A materially positive final increment implies a rising lower-bound estimate; a small terminal change or non-monotone oscillation within sampling noise is treated as convergence.
  • Evaluation: Class-normalized evaluation is required for cohorts that are not artificially balanced.The weighting implements equal-prior expectations despite the cohort prevalence.

Real-Cohort Experimental Details

The three cohort audits use distinct measurement channels and validation designs, with diagnostics distinguishing converged frontiers from lower bounds. Results show near-frontier learners in some settings, strong complementarity across channels, and large learner gaps for deficient models.

  • UCI Diabetes 130-US Hospitals Readmission: UCI contains 99,343 encounters from 69,990 patients, with patient-grouped folds and patient-level bootstrap resampling.The cohort excludes death and hospice discharges and has 30-day-readmission prevalence 0.1139.
  • Convergence Diagnostics: BRFSS stabilizes across training fractions, while NHANES oscillates non-monotonically and is treated as converged rather than rising.The BRFSS final increment is −0.0002; the NHANES sequence drops 0.0037 from 0.5 to 0.75 and is treated as finite-sample oscillation.
  • UCI Diabetes 130-US Hospitals Readmission: UCI’s ceiling is 0.6225 [0.6213, 0.6240], achieved balanced accuracy is 0.6223, and G = +0.0002.The underfit sequence ends with a +0.0029 increment, so the frontier is reported as a lower bound.
  • BRFSS and NHANES Channels: BRFSS uses 253,680 telephone-survey respondents and separates perception channel A from recalled-diagnosis channel B.Questionnaire, measured, and joint channel definitions are also specified for NHANES, excluding glycemic outcome variables to prevent leakage.
  • NHANES HbA1c: NHANES questionnaire and measured ceilings differ by +0.0042 with overlapping intervals, whereas the joint ceiling is 0.7623 and complementarity is +0.0471.The joint interval is disjoint from the measured interval, so the gain reflects complementary decision information rather than marginal dominance.
  • Complete Learner Panels: The NHANES learner panel reports gaps of +0.0035, +0.0196, and +0.0081 for LR, RF, and GBDT, while MLP has AUROC 0.8385 and gap +0.1933.The regularized posterior learner is HistGradientBoosting with early stopping and native NaN handling.

Complete Learner Panels

The learner panels compare frontier estimates with achieved balanced accuracy across model families, showing that similar estimated frontiers can coexist with sharply different learner performance.

  • UCI Learner Panel: Table 4 reports the UCI learner panel, with gap defined as ceiling minus achieved balanced accuracy.This decomposition distinguishes the fixed-channel frontier from learner-specific performance.
  • BRFSS Learner Panel: Table 5 reports the corresponding BRFSS learner panel.
  • NHANES Joint-Channel Learner Panel: Table 6 reports the NHANES joint-channel learner panel.The panel supports comparison of learner gaps for the joint measurement channel.

Clinical Evidence-Synthesis Protocol

The evidence synthesis uses PRISMA-guided screening and preserves heterogeneous clinical metrics rather than pooling them into a universal estimand. It analyzes 104 task-level observations from 30 source publications across more than 18 categories.

  • Scope and Rationale: The synthesis is framed as descriptive external context, not as a pooled estimate of a universal ceiling.Heterogeneous outcomes, validation schemes, populations, and metrics are retained on their original scales.
  • Search Protocol: Searches followed PRISMA 2020 and umbrella-review guidance across PubMed, PubMed Central, ScienceDirect, Springer-Link, Authorea, and arXiv.
  • Eligibility: Eligible reports included English-language systematic, scoping, or meta-analytic reviews and large primary studies with n ≥500 using structured clinical or patient-reported inputs.Studies had to report a quantitative predictive metric; imaging-only reports were excluded.
  • Screening and Dataset: Screening yielded 1,117 database records and 19 manually identified records, while the analytic dataset contains 30 source publications and 104 task-level observations.Figure 8 distinguishes source publications from extracted task observations rather than treating them as one denominator.
  • Data Extraction: Extracted fields included disease category, learner family, sample size, predictive metrics, validation design, class-specific recall when available, and multimodal status.AUROC, balanced accuracy, raw accuracy, F1, and AUPRC were not pooled as one estimand.

Result Analysis

The synthesis finds recurring middle-range performance across many structured-clinical categories, with diminishing gains from increasingly complex learners and higher performance when measurement channels change. These patterns are descriptive and require qualification because tasks, channels, outcomes, cohorts, metrics, and validation designs vary.

  • Interpretive limits: Reported category ranges are not pooled effect estimates or directly comparable channel frontiers because disease categories combine heterogeneous tasks, outcomes, metrics, channels, cohorts, and validation designs.Examples include ICU/sepsis spanning 0.75 to 0.99, breast cancer 0.57 to 0.97, and autoimmune/rheumatology 0.63 to 0.92; small studies also risk unstable or overoptimistic estimates.
  • Cross-category patterns: 104 task-level observations across more than 18 disease categories show a recurrent middle performance region across diverse clinical domains.The pattern appears across orthopedic surgery, cardiovascular disease, obstetrics, endocrine/renal disease, stroke, mental health, and several oncology tasks.
  • Model-family patterns: Logistic regression typically reports 0.78–0.80, SVM 0.80–0.82, random forests 0.81–0.84, and XGBoost or gradient boosting 0.83–0.87.Multilayer perceptrons and tabular deep-learning systems generally add little beyond boosting, with typical values around 0.84–0.88.
  • Model-family patterns: Within fixed cohorts, estimated frontiers remain comparatively stable while achieved balanced accuracy can differ dramatically across architectures.In BRFSS and NHANES, MLP AUROC remains competitive while thresholded balanced accuracy collapses, indicating that ranking performance need not yield a useful decision rule.
  • Measurement-channel comparisons: Clinical-only results lie roughly between 0.75 and 0.83, whereas corresponding multimodal results lie roughly between 0.83 and 0.90 in six selected comparisons.These contrasts illustrate possible frontier movement when an added channel contributes decision-relevant information, but they are not harmonized within-cohort modality ablations.
  • Measurement-channel comparisons: NHANES questionnaire and measured marginal frontiers are statistically indistinguishable, while the joint frontier rises by +0.0471 over the better marginal.The result supports complementarity rather than a universal claim that measured data intrinsically dominate questionnaires.
Loading 2609.01909v1…