Source-linked AI summary

Separating Voice from Age in COPD Screening

George P. Kafentzis, Nikoletta Arvaniti

arXiv:2608.21599v1eess.AScs.LGcs.SDeess.SP

TL;DR

Voice-based COPD screening may confuse disease with age because both COPD and voice change with age. The paper re-evaluates a public sustained-phonation corpus using participant-level, repeatedly age-matched cohorts and raw confounder controls. Acoustic models excluding age retain discrimination while age-containing models perform worse, supporting measurable non-age acoustic signal but not clinical viability.

  • Problem

    COPD and voice are both age-associated, so screening performance may reflect age rather than disease; standard age removal does not eliminate acoustic age proxies.

  • Method

    The study applies a strictly participant-level protocol with repeated one-to-one age matching, raw age and gender controls, paired bootstrap resampling, and replication under additional learners.

  • Results

    ROC-AUC 0.665–0.717 for age-excluding configurations versus 0.531–0.679 for configurations containing age, with raw age and gender at chance on matched cohorts.

  • Takeaways & Limitations

    A non-age acoustic signal is present and measurable in these data, while standard evaluation cannot distinguish it from confounding.

  • Takeaways & Limitations

    The analysis uses approximately 24 matched pairs with effective independent units of 24, producing wide intervals, and generalization beyond one site, language, vowel, and 30 cases is untested.

Abstract

from arXiv · show

Voice has been proposed as a low-cost screening signal for chronic obstructive pulmonary disease (COPD). COPD is strongly age-associated and voice changes with age, thus such results admit a trivial alternative explanation. We re-evaluate a public sustained-phonation corpus ($1246$ recordings, $68$ participants) under a strictly participant-level protocol. We therefore evaluate on repeatedly drawn age-matched cohorts and report the discrimination achieved by the confounders themselves on those same cohorts. Where raw (unmodelled) age ($0.510$ $[0.469, 0.551]$) and raw gender ($0.479$) are both measured at chance, acoustic models excluding age retain ROC-AUC $0.717$ $[0.552, 0.859]$ and average precision $0.747$ $[0.581, 0.892]$ against a one-to-one baseline of $0.5$, whereas models containing age fall to $0.531$--$0.679$. The separation is reproduced by two further learners with fixed hyperparameters. Two findings have broader methodological implications: models trained with age transfer less effectively to an age-balanced target cohort than otherwise identical models trained without age, and fourteen classical voice-quality and perturbation measures achieve comparable discrimination to a $55$-dimensional combined representation. We conclude that a non-age acoustic signal is present, that confounding by recording conditions cannot be excluded from the released features, and that the evaluation protocol in standard use cannot distinguish these possibilities.

I. INTRODUCTION

COPD voice-screening results may reflect age rather than disease because COPD and voice are both age-associated. This paper therefore audits confounding and evaluates whether acoustic discrimination survives verified age matching.

  • COPD is often asymptomatic early, while spirometry requires clinic attendance, trained personnel, and patient effort, motivating low-cost screening signals.
  • Age is a plausible alternative explanation because COPD is age-associated and voice changes through vocal-fold, respiratory, and acoustic mechanisms.
  • Removing age from predictors is insufficient when acoustic features retain redundant age information that models can exploit as proxies.
  • The study re-evaluates 1246 recordings from 68 participants using 107 predictors spanning demographics, self-reported conditions, static acoustics, and MFCC-derived features.
  • The paper audits claimed matching, tests confounder-only and acoustic models under participant-level age matching, and examines whether standard pooled evaluation can detect non-age signal.
  • The authors limit their claim to measurable non-age acoustic signal in these data, not clinical viability, given roughly two dozen pairs from one site, language, and recording protocol.

B. Within-Patient Longitudinal Monitoring

Within-patient longitudinal monitoring addresses a different question from cross-sectional screening: whether a patient’s condition changes over time. Repeated measurements can hold stable speaker characteristics fixed, but study designs still vary in validity and scope.

  • B. Within-Patient Longitudinal Monitoring: Within-patient monitoring uses each patient as their own control, largely removing between-subject confounding from age, sex, anatomy, and habitual voice quality.
  • B. Within-Patient Longitudinal Monitoring: Nallanthighal et al. detect COPD exacerbation from speech in 40 patients, reporting 75.12 % accuracy and 0.85 sensitivity.
  • B. Within-Patient Longitudinal Monitoring: The breathing-focused model estimates a physiologically meaningful latent breathing rate, but the all-COPD cohort supports monitoring rather than screening.
  • B. Within-Patient Longitudinal Monitoring: TACTICAS collected daily home recordings from 73 participants and captured 38 exacerbations from 35 participants, illustrating deployment conditions and the importance of adherence measurement.
  • B. Within-Patient Longitudinal Monitoring: Some longitudinal evaluations distribute recordings from the same participants across folds, allowing reported performance to include an unknown speaker-recognition component.
  • B. Within-Patient Longitudinal Monitoring: A study predicting COPD Assessment Test deviation reports R2 = 0.972, r = 0.998, and 94 % accuracy, but its nine participants and non-speaker-disjoint folds warrant caution.

C. Speaker Identity and Confound Control

Speaker identity and demographic confounding require explicit controls because representations can encode nuisance information. This paper complements representation-level methods by changing the evaluation so confounder discrimination is directly measured.

  • C. Speaker Identity and Confound Control: Speaker-disentangled representations suppress speaker information while optimizing respiratory-status classification, improving stable-versus-exacerbated AUC from 0.897 to 0.910.
  • C. Speaker Identity and Confound Control: The present work instead constructs cohorts where the confounder is verifiably uninformative and reports the confounder’s own discrimination as a control.
  • C. Speaker Identity and Confound Control: Prior cross-sectional screening studies remain exposed to between-subject confounding, whereas longitudinal designs largely avoid it through within-patient comparisons.
  • C. Speaker Identity and Confound Control: The authors identify a gap in verifying demographic balance, testing confounder-only discrimination, and distinguishing learned signal from confounder predictions.
  • C. Speaker Identity and Confound Control: The corpus contains 1246 sustained-vowel recordings from 68 participants and 107 predictors, with recording counts severely unbalanced across participants.
  • C. Speaker Identity and Confound Control: Figure 1 shows that age imbalance depends on weighting: participant-level and recording-level group differences reverse in direction.

B. The Age Imbalance

Participant-level analysis reveals substantial age imbalance despite the source publication’s reported matching, with the cohort’s age structure allowing age alone to discriminate COPD cases from controls.

  • Participant-level imbalance: 0.726 standardized mean difference shows large participant-level age imbalance, despite reported group mean ages differing by less than one month.COPD participants averaged 72.73 years, while controls averaged 64.68 years.
  • Participant-level imbalance: 7.7 years separates participant-weighted group means, whereas recording-weighted means reverse the apparent difference to 2.1 years in the opposite direction.Unequal recording counts and opposing within-group count–age correlations drive the reversal.
  • Participant-level imbalance: Recording-level summaries can conceal—and even invert—cohort imbalance when repeated-measures datasets assign unequal numbers of recordings to participants.Balance must therefore be reported at the classifier’s unit of analysis.
  • Age-only discrimination: 0.721 ROC-AUC is achieved by the non-monotone score −|age −78|, compared with 0.677 for a monotone age score.The non-monotone score wins 800 of 1110 COPD–control pairs, giving learners access to substantial age-only discrimination.
  • Age-only discrimination: Pooled evaluation cannot distinguish acoustic learning from exploitation of the recruitment age distribution.The cohort’s prevalence peaks at ages 70–75 and declines because COPD participants are bounded above at 81 years.

D. Acoustic Features as Age Proxies

Acoustic descriptors contain limited age information, and the strongest marginal COPD discriminators are largely distinct from the most age-correlated features; the evaluation is designed to isolate these effects at participant level.

  • Age association: Median absolute age correlation is 0.19 and the maximum is 0.45 across 49 static acoustic descriptors.Seven descriptors exceed 0.3, but none reaches 0.5.
  • Age association: MFCC standard deviations are the most age-correlated features, with |ρ| between 0.36 and 0.45, yet they discriminate comparatively weakly.The figure places these features among the strongest age associations but weaker COPD discriminators.
  • Discrimination versus age correlation: 0.733 is the highest listed univariate COPD AUC, attained by localJitter, while the strongest perturbation measures have |ρ| between 0.16 and 0.24 with age.The listed high-AUC measures include localJitter, ppq5Jitter, sd_F0_list, ddpJitter, and rapJitter.
  • Other candidate confounders: ROC-AUC 0.490 for raw gender indicates no gender confounding, while recording count is likewise uninformative at ROC-AUC 0.545.Gender proportions are similar between COPD participants and controls.
  • Other candidate confounders: Phonation duration reaches ROC-AUC 0.602, but its interpretation is unresolved because the released features do not identify the recording protocol.The analysis therefore removes duration from the acoustics-only configuration.
  • Discrimination versus age correlation: Univariate AUCs are in-sample and order features rather than estimate generalization performance.Fig. 3 compares age association with univariate COPD discrimination at participant level.

B. Participant-Level Cross-Validation

The evaluation uses grouped, repeated cross-validation and participant-level weighting, aggregation, and scoring so that repeated recordings cannot substitute for independent participants.

  • Participant-level design: 25 outer partitions combine five-fold stratified grouped cross-validation with five random seeds, keeping every participant within one split.The same partitions are reused across all twelve feature configurations.
  • Participant-level design: Equal total training weight is assigned to each participant despite unequal recording counts, while evaluation remains unweighted after participant-level aggregation.Recording-level weights are applied only to the training objective.
  • Score aggregation: One participant-level score is produced by aggregating each participant’s recording scores in the logit domain.The per-recording probabilities are clipped before logit averaging.
  • Score aggregation: Arithmetic averaging of scores does not alter the ordering of configurations, despite logit averaging’s sensitivity to confident individual predictions.Logit averaging is treated as pooling repeated measurements of one underlying state.

4) Pooled scores:

The evaluation uses participant-level resampling and repeatedly drawn one-to-one age-matched cohorts, with confounder discrimination verified on each evaluation cohort. Hyperparameter selection favors simpler models within one standard error of the best inner-fold AUC.

  • Model selection: The one-standard-error rule selects a simpler hyperparameter configuration whose inner-fold mean AUC is within one standard error of the maximum.Complexity is ordered by depth, iterations, stronger L2 regularization, and learning rate.
  • Model selection: Inner-fold AUC has an estimated standard error of roughly 0.10, and maximizer instability moved pooled ROC-AUC by up to 0.08 for one- and two-predictor configurations.Configurations with 50 or more predictors reproduced within 0.02 across otherwise identical executions.
  • Participant-level evaluation: 2000 bootstrap resamples of 67 pooled participant scores produce percentile intervals for ROC-AUC and average precision.Average precision is reported with its class-prevalence chance baseline, while configuration comparisons use 20,000 paired, class-stratified bootstrap resamples.
  • Matched-cohort evaluation: Each of 2000 repetitions draws a one-to-one age-matched cohort by pairing participants within a two-year caliper, yielding approximately 24 pairs on average.Matching is applied at evaluation time to models trained on the unmatched cohort.
  • Matched-cohort evaluation: Raw age and gender are scored on every matched cohort, and only cohorts where raw age discriminates at chance support an age-controlled interpretation.This verification uses the confounders’ own AUC rather than relying only on a difference-of-means balance statistic.

F. Feature Importance and Uncertainty Quantification

The paper preserves repeated-measures structure when estimating feature importance and describes conformal prediction behavior at the participant level. It also repeats the protocol with two fixed-hyperparameter learners to test whether findings depend on the original learner or selection procedure.

  • Feature Importance: Permutation importance is computed at participant level by swapping features through donor participants while preserving within-participant recording structure.Importance is the average drop in subject-level AUC across five repetitions per outer split.
  • Uncertainty Quantification: Conformal prediction uses participant-level aggregated scores, label-stratified grouped calibration folds, and a pooled threshold for marginal coverage.The threshold is formed from all calibration scores rather than from separate class-conditional thresholds.
  • Uncertainty Quantification: Coverage is reported with abstention rate because prediction sets may be empty, singleton, or contain both labels.A set containing both labels covers the truth trivially, making coverage alone uninformative.
  • Learner robustness: The complete protocol is repeated with histogram-based gradient boosting and L2-regularized logistic regression using fixed hyperparameters and no inner cross-validation.This removes selection instability while retaining the matched evaluation framework.
  • Learner robustness: The paper presents pooled analysis before matched-cohort analysis, then tests whether the findings depend on learner choice.The ordering treats pooled evaluation’s inability to answer the target question as a result in itself.

A. Pooled Evaluation

Pooled evaluation cannot separate disease-related acoustic signal from age-related confounding: age alone can match model performance, and pooled contrasts are inconclusive. After age matching, age-free models remain comparatively stable while age-containing models lose performance.

  • Pooled discrimination: 0.671–0.741 ROC-AUC spans the pooled configurations, making a one-predictor age model (0.709) statistically indistinguishable from a 103-predictor combined model (0.741).The interval range is narrower than the uncertainty intervals, so configuration ordering is not interpretable.
  • Pooled discrimination: 0.721 ROC-AUC from a hand-specified non-monotone age function shows that pooled performance can be achieved without acoustic information.Pooled evaluation therefore cannot distinguish learned respiratory pathology from the recruitment age distribution.
  • Pooled comparisons: None of sixteen paired pooled comparisons excludes zero, while adding symptom variables changes ROC-AUC by +0.001 [−0.040, +0.037].Removing phonation duration changes ROC-AUC by +0.005 [−0.040, +0.049], and gender changes it by +0.005 [−0.068, +0.087].
  • Pooled comparisons: −0.026 [−0.151, +0.095] and +0.010 [−0.098, +0.117] are the reported incremental values of age, but the pooled design cannot support a conclusion about them.Acoustics versus demographics alone ranges from −0.045 to +0.025 with interval widths of 0.26–0.42.
  • Matched-cohort evaluation: 0.510 [0.469, 0.551] raw-age ROC-AUC and 0.479 raw-gender ROC-AUC are at chance on matched cohorts.The matched-cohort age SMD is 0.026, down from 0.726 in the unmatched cohort.
  • Matched-cohort evaluation: −0.008 ROC-AUC is the average matching shift for age-free configurations, versus −0.108 for age-containing configurations.The age-free models remain stable under repeatedly drawn age-matched subsets.
  • Matched-cohort evaluation: 0.755 [0.612, 0.873] PR-AUC is achieved by fourteen a priori clinical measures, while AC achieves 0.747 [0.581, 0.892].Matched prevalence is exactly 0.5 in every replicate under paired resampling.
  • Matched-cohort evaluation: Only AC excludes zero against raw age, at +0.207 [+0.027, +0.365], across twelve matched-cohort contrasts.The authors characterize this single uncorrected exclusion as suggestive and no more.

C. Independent Replication

The matched-cohort separation persists across histogram-based gradient boosting and logistic regression, indicating that the age-free versus age-containing contrast is not specific to CatBoost. The linear results also expose limits of pooled evaluation and support a non-linear acoustic–disease relationship, with regularization as a caveat.

  • Histogram-based boosting: 0.702 versus 0.556 mean matched ROC-AUC separates age-free from age-containing configurations under histogram-based boosting.Under CatBoost, the corresponding means are 0.685 and 0.603.
  • Histogram-based boosting: 0.902 Pearson r and 0.935 Spearman ρ show strong agreement between CatBoost and histogram boosting across configurations.Their mean absolute difference is 0.035.
  • Histogram-based boosting: All five age-free histogram-boosting configurations clear the ROC bar, while none of seven age-containing configurations does.The highest PR-AUC lower bound in the study is PTRB at 0.779 [0.651, 0.888].
  • Logistic regression: 0.497 versus 0.583 reverses the pooled logistic-regression ordering of age versus acoustics on matched cohorts.The same dissociation appears under both boosted ensembles.
  • Cross-learner comparison: +0.082, +0.147, and +0.066 are the age-free-minus-age-containing family-mean differences for CatBoost, histogram boosting, and logistic regression.The linear model is underpowered: no configuration clears chance and its intervals are approximately 0.37 wide.
  • Logistic regression: −0.043 [−0.091, −0.005], −0.077 [−0.150, −0.015], and −0.110 [−0.222, −0.010] are the only logistic paired contrasts excluding zero against raw age.They correspond to the three age-dominant configurations A, D, and DH.
  • Interpretation: The acoustic–disease relationship may not be well approximated linearly, but the learner comparison also confounds model family with fixed rather than tuned regularization strength.This limits interpretation of the gap between linear and boosted learners.

D. Feature Importance and Conformal Prediction

Feature-importance analysis highlights F0 variability in age-free models, while age dominates when explicitly supplied. Conformal prediction shows that nominal coverage requires frequent abstention, limiting screening decisiveness.

  • Feature Importance: F0 standard deviation ranks first in both analyzed age-free configurations, with mean AUC drops of 0.040 (DAC-A) and 0.071 (PTRB).
  • Feature Importance: 94 % of total importance in the 50-predictor age-free configuration is attributed to F0 variability, compared with 60 % in the 14-predictor clinical set.
  • Feature Importance: Age ranks first in ALL with importance 0.060, more than double the highest-ranked acoustic feature.
  • Feature Importance: Importance rankings are unstable across resamplings, with mean Spearman correlations of 0.032, 0.111, and 0.202 for the 50-, 55-, and 14-predictor configurations.The authors therefore do not warrant claims about relative individual-feature importance, despite F0 variability leading across three feature sets.
  • Conformal Prediction: At α = 0.10, nominal 90 % coverage requires abstention on 61–79 % of participants, while α = 0.20 reduces abstention to 26–57 %.Observed coverage met or exceeded nominal levels, but the exploratory construction carries no finite-sample guarantee.
  • Conformal Prediction: At α = 0.20, age-free configurations achieve decisive-and-correct predictions for 0.487–0.554 of all participants, versus 0.305–0.340 for demographic-only configurations.This product combines decision frequency and correctness, which should be interpreted jointly rather than separately.
  • Matched Evaluation: Age-excluding acoustic models retain ROC-AUC 0.665–0.717 and average precision 0.734–0.755 on matched cohorts, whereas age-containing models reach 0.531–0.679.Raw age and gender discriminate at chance on these cohorts against a one-to-one baseline of 0.5.

A. Implications for Cross-Sectional Screening Studies

The paper argues that cross-sectional voice-screening results require participant-level balance checks and direct confounder controls. It also identifies recording conditions as the principal unresolved alternative explanation and sets boundaries on generalization.

  • A. Implications for Cross-Sectional Screening Studies: Repeated-measures speech datasets can invert apparent group differences when unequal recording contributions are summarized at the recording rather than participant level.In this corpus, a single control contributing a quarter of recordings reversed the sign of the group age difference.
  • A. Implications for Cross-Sectional Screening Studies: Participant-level cohort balance and the confounder’s discrimination on the evaluation cohort should both be reported, because mean-balance statistics and rank-based classifier metrics can disagree.
  • A. Implications for Cross-Sectional Screening Studies: Cross-sectional accuracies of 75–95 % remain measurements, but their division between pathology and demographics is unclear when age differs between cases and controls.Within-patient longitudinal designs are described as largely immune to this between-subject confounding.
  • A. Implications for Cross-Sectional Screening Studies: Supplying age during training reduces transfer to age-balanced cohorts, with matching-related performance drops shrinking as more acoustic predictors accompany the demographic block.Reported drops range from −0.178 to −0.062 along the DAC ladder and from −0.178 to −0.085 along the ALL ladder.
  • A. Implications for Cross-Sectional Screening Studies: Including a confounded demographic covariate can cause models to under-learn acoustic structure, so excluding it at training time is a modeling decision rather than merely a post-hoc robustness check.
  • A. Implications for Cross-Sectional Screening Studies: Recording conditions remain the leading non-pathological explanation for the residual signal, but the released features cannot exclude device, microphone, environment, gain, session, or operator confounding.Such confounding would be unaffected by age matching and would appear in spectral and perturbation descriptors.
  • A. Implications for Cross-Sectional Screening Studies: The present result supports a non-age, non-gender acoustic signal while leaving respiratory versus instrumental origin unresolved.The authors recommend releasing recording metadata for cross-sectional benchmark corpora.
  • A. Implications for Cross-Sectional Screening Studies: The matched analysis averages 24.0 pairs, leaving wide intervals and making matched and pooled estimates different estimands rather than corrected versions of one another.The effective number of independent units is 24 because bootstrap resampling uses matched pairs.

VII. CONCLUSIONS

The paper re-evaluates a public COPD voice corpus using participant-level, age-matched evaluation and direct confounder controls. Age-free acoustic models retain discrimination, but recording-condition confounding and clinical generalization remain unresolved.

  • VII. CONCLUSIONS: The corpus has a participant-level age standardized mean difference of 0.726 despite being described as age-matched within five years.Eight of 37 controls have no case within five years, and recording-level weighting inverts the apparent group age difference.
  • VII. CONCLUSIONS: Under pooled imbalance, all twelve feature configurations clear chance, while a hand-specified non-monotone age function reaches ROC-AUC 0.721 and none of sixteen paired comparisons excludes zero.
  • VII. CONCLUSIONS: Under participant-level age-matched evaluation, age-free acoustic models retain ROC-AUC 0.717 [0.552, 0.859] and average precision 0.747 [0.581, 0.892].Raw age and gender discriminate at chance, whereas models containing age fall toward chance.
  • VII. CONCLUSIONS: Training with a confounded covariate reduces transfer to age-balanced evaluation cohorts, making demographic exclusion a modeling decision rather than only a robustness check.Fourteen classical voice-quality and perturbation measures match the discrimination of a 55-dimensional combined representation.
  • VII. CONCLUSIONS: A non-age, non-gender acoustic signal is measurable, but its respiratory or instrumental origin cannot be settled without unreleased recording-condition metadata.The standard cross-sectional evaluation protocol cannot distinguish these possibilities.
  • VII. CONCLUSIONS: Reporting participant-level balance and confounder discrimination on the evaluation cohort would materially improve interpretability of the literature.
Loading 2608.21599v1…