Source-linked AI summary

ReliaGate: Reliability Routing for Low-Stakes Wearable Stress Prediction

Jaden Moon, Yu Wu, Arvind Pillai, Andrew Campbell

arXiv:2608.15951v1cs.LGcs.HC

TL;DR

Wearable stress systems need to decide when an individual prediction is reliable enough to surface, because aggregate accuracy does not resolve per-output availability. ReliaGate post-hoc routes unchanged classifier labels using reliability cues; evidence was favorable but not uniform across four datasets, with UBFC-Phys showing significant coverage and accepted-risk advantages while E4 outcomes were mixed.

  • Problem

    Aggregate accuracy does not determine whether an individual wearable stress label is reliable enough to surface amid artifacts, participant variation, and dataset shift.

  • Method

    ReliaGate uses a post-hoc gate to surface or withhold unchanged locked-classifier labels based on confidence, signal quality, agreement, atypicality, and train-fitted geometry cues.

  • Results

    Across four evaluations, evidence favored geometry-augmented routing but was not uniform; on UBFC-Phys, coverage improved by +0.101 and accepted risk decreased by −0.056, with both intervals excluding zero.

  • Takeaways & Limitations

    ReliaGate is a framework for studying surfaced-label error, output availability, and uneven subject-level access rather than a universal geometry improvement.

  • Takeaways & Limitations

    Protocol labels are laboratory proxies rather than clinical or free-living stress, and the method does not provide clinical or finite-sample risk guarantees.

Abstract

from arXiv · show

We study when a wearable stress system should surface a prediction rather than change it. In low-stakes reflection and summary settings, aggregate accuracy is insufficient because withholding can reduce error while leaving some people with little or no information. We formulate fixed-label reliability routing: after a locked classifier emits a protocol-defined stress/non-stress label, a post-hoc gate surfaces that unchanged label or withholds it as unavailable. ReliaGate assembles established confidence, signal-quality/trust, agreement, train-standardized atypicality, and train-fitted geometry cues into a post-hoc correctness score. We evaluate four wearable datasets using subject-disjoint folds, validation-selected routing, paired held-out-subject intervals, and pooled and per-subject analyses. WESAD point estimates favored ReliaGate, UBFC-Phys primary coverage/risk intervals favored ReliaGate, and E4 checks were mixed. ReliaGate provides an operational framework for studying surfaced-label error, output availability, and accepted-output distribution across subjects, without revising labels or providing clinical or finite-sample risk guarantees.

1 Introduction

The paper frames wearable stress prediction as a reliability-routing problem: a post-hoc gate decides whether to surface a locked classifier’s unchanged label or withhold it as unavailable. ReliaGate combines multiple established reliability cues and evaluates routing across four wearable datasets with subject-disjoint, fold-local procedures and per-subject availability analyses.

  • Motivation: Wearable stress systems are positioned for low-stakes reflection and summaries, using signals such as BVP, EDA, and TEMP with protocol-defined stress/non-stress labels.The introduction distinguishes these applications from diagnosis.
  • Motivation: The key question is whether each classifier output is reliable enough to show, because artifacts, participant variation, protocol differences, and dataset shift undermine aggregate accuracy.Confidence calibration can also degrade under dataset shift.
  • Reliability routing: Reliability routing applies a post-hoc gate after a locked classifier, surfacing its unchanged label or withholding it as unavailable.Coverage measures the fraction surfaced, while accepted risk measures errors among surfaced labels; both are needed because withholding can lower risk.
  • ReliaGate: ReliaGate scores correctness using confidence, signal quality and modality trust, cross-modal agreement, train-standardized atypicality, and train-fitted embedding geometry.The scorer is designed as a leakage-audited, post-hoc assembly without modifying the upstream classifier.
  • Evaluation: The evaluation uses four wearable datasets, subject-disjoint folds, fold-local scorer and threshold selection, paired held-out-subject intervals, and pooled and per-subject analyses.Per-subject analysis tests whether routing leaves held-out subjects with little or no output.

2 Related Work

Wearable stress systems must assess whether a fixed prediction is reliable enough to show because performance varies across participants, sensors, protocols, stressors, and signal quality. ReliaGate builds on selective-classification ideas by integrating established reliability cues with validation-locked routing and per-subject availability analysis.

  • Wearable stress sensing and reliability: Wearable stress sensing uses PPG-derived BVP, EDA, skin temperature, and activity signals, whose prediction performance can vary across participants, sensors, protocols, and stressors.PPG is also vulnerable to motion artifacts, creating additional signal-quality variation.
  • Wearable stress sensing and reliability: The central well-being-system challenge is deciding whether a particular stress label is reliable enough to show, not only predicting the label.This frames surfacing as a reliability decision after prediction.
  • Selective classification and ReliaGate: ReliaGate scores likely correctness after the wearable label is fixed, then surfaces or withholds that unchanged label.The approach is fixed-label surfacing rather than label revision.
  • Selective classification and ReliaGate: ReliaGate integrates established cues using leakage-audited, validation-locked routing and adds a per-subject availability lens.The contribution is presented as an integration and evaluation framework, not a new calibration method, distance metric, or OOD detector.

3 ReliaGate: Representation-Augmented Reliability Routing

ReliaGate is a post-hoc gate that routes a locked stress/non-stress label unchanged or withholds it as unavailable, using multimodal reliability evidence and representation geometry. It learns correctness for ranking and thresholding, while validation-selected routing and shared-subject geometry impose important methodological limits.

  • Routing framework: ReliaGate scores a locked classifier’s fixed label and either surfaces it unchanged or withholds it as unavailable.The reference label is used only afterward for evaluation.
  • Reliability evidence: The score combines confidence, signal quality/trust, cross-modality agreement, train-standardized atypicality, and train-fitted geometry.Geometry references are fit using outer-training data, while the canonical feature-model predictions are subject-wise out of fold.
  • Correctness modeling: ReliaGate learns binary correctness rather than stress class, using post-hoc correctness candidates and thresholding within reject-option or selective-classification frameworks.Neither the reference label nor correctness target is available at inference time.
  • Feature construction: Train-standardized atypicality cues are heuristic rather than formal shift estimates, with feature manifests excluding labels, targets, identifiers, thresholds, and routing outcomes.The non-geometry vector includes modality quality/trust, cross-modality summaries, and atypicality norms.
  • Methodological limitations: Primary geometry is not subject-excluded because selector-training references retain other query-subject windows and the encoder may have seen that subject.A WESAD sensitivity removes the query subject from references but retains encoder exposure, testing reference sharing rather than subject-excluded training.
  • Methodological limitations: Validation selects the scorer family and threshold, but overlapping validation windows can make selection optimistic; the routing criterion is empirical, not a finite-sample guarantee.System-level comparison is primary because NG and RG may select different scorer families.

4 Task, Datasets, and Evaluation Protocol

The evaluation uses four independently trained wearable datasets for within-dataset, subject-held-out analyses, with protocol-defined binary labels and fixed 60-second windows. It compares no-withholding and post-hoc routing systems using pooled and per-subject availability, error, and performance metrics with paired held-out-subject intervals.

  • Datasets and folds: Four datasets provide complementary within-dataset evaluations, with each outer fold holding out one analyzed test subject.WESAD supports primary mechanism and robustness analyses, UBFC-Phys provides controlled BVP/EDA evaluation, and E4 datasets provide protocol checks.
  • Tasks and preprocessing: Protocol-defined binary tasks map dataset periods to stress or non-stress labels using 60 s windows with 10 s strides after specified trims.WESAD excludes other codes and unused ACC; UBFC-Phys maps T1 to non-stress and T2–T3 to stress; EmpaticaE4Stress maps rest/task blocks.
  • Routing systems: All systems act on the same fixed label: Always accepts every label, Confidence-only uses predicted-class confidence, NG adds quality, trust, agreement, and atypicality, and RG adds geometry.NG and RG select logistic or GBDT correctness scorers within each fold; the primary comparison is system-level.
  • Evaluation metrics: Reported metrics jointly quantify accepted-window coverage, conditional error risk, accepted-and-wrong fraction, and subject reach, alongside accepted macro-F1 and balanced accuracy.SF = Cov × Risk when outputs exist, and pooled values weight subjects by accepted-window counts; per-subject analyses describe output distribution.
  • Evaluation protocol: Paired 95% bootstrap percentile intervals use 10,000 shared held-out-subject resamples, while test data remain reporting-only and selection uses validation within each fold.Training and selection are not rerun for the intervals; validation/test geometry uses only outer-training references.

5 Results

ReliaGate generally improved coverage and accepted-risk point estimates over no-geometry routing on WESAD and UBFC-Phys, while E4 results were mixed. Subject-level outcomes and geometry analyses supported gains but also showed trade-offs and uncertainty.

  • WESAD: On WESAD, RG accepted 1,908 windows with 334 errors versus NG’s 1,594 windows with 371 errors, while both reached 14 of 15 subjects.RG therefore surfaced 314 more accepted windows and made 37 fewer accepted errors on different subsets.
  • WESAD: Across four α values and six test offsets, RG had higher-coverage and lower-risk point estimates than NG on WESAD.Geometry permutations increased mean AURC by 0.115 (95% interval [0.068,0.168]) in all 15 folds, consistent with ranking sensitivity rather than a causal effect.
  • WESAD: WESAD subject-level results favored RG for median coverage, 0.749 versus 0.601, and 25%-coverage reach, 13/15 versus 11/15, but NG had lower median accepted risk, 0.107 versus 0.163.Alternative geometry references and a post-hoc GBDT comparison also reported RG coverage/risk of 0.654/0.177 and 0.679/0.159, respectively.
  • UBFC-Phys: On UBFC-Phys, RG accepted 1,244 windows with 318 errors versus NG’s 1,057 windows with 329 errors, with coverage difference +0.101 and accepted-risk difference −0.056.Both paired 95% intervals excluded zero in the favorable direction.
  • UBFC-Phys: UBFC-Phys geometry permutation increased mean AURC by 0.058, while median subject coverage rose to 0.758 from 0.667 and zero-coverage subjects fell from six to three.The AURC change was positive in 45 of 56 folds.
  • E4: E4 outcomes were mixed: RG improved coverage, mF1, and BAcc on both datasets, but risk, silent-failure rate, Reach, and paired intervals varied by dataset.On EmpaticaE4Stress, RG had lower accepted risk and unchanged Reach but higher SF; on PhysioNetE4, RG had higher risk and SF and lower Reach, with both datasets’ paired intervals including zero.

6 Discussion, Limitations, and Ethics

Selective wearable systems cannot be judged by accepted risk alone because surfaced-label error and output availability can differ across people. ReliaGate is supported as a framework for studying coverage, surfaced-label error, and subject-level access, not as a universal geometry improvement.

  • Interpretation and implications: Selective wearable systems cannot be judged by accepted risk alone.
  • Interpretation and implications: Pooled coverage and risk favored ReliaGate, but median subject risk rose despite higher median coverage.Surfaced-label error and output availability can move differently across people.
  • Interpretation and implications: Only UBFC-Phys had favorable primary coverage/risk intervals excluding zero, while E4 outcomes were mixed.
  • Interpretation and implications: The evidence supports ReliaGate as a framework for studying coverage, surfaced-label error, and subject-level access, not as a universal geometry improvement.

7 Conclusion

ReliaGate treats low-stakes wearable stress inference as fixed-label routing, deciding whether to surface an unchanged stress/non-stress prediction. Across four within-dataset evaluations, it jointly considers output availability and surfaced-label error, with favorable but non-uniform evidence for geometry-augmented routing.

  • Fixed-label routing: ReliaGate uses a separate gate to surface or withhold an unchanged protocol-derived stress/non-stress label.This frames the system as fixed-label routing rather than label revision.
  • Evaluation framework: ReliaGate combines wearable reliability cues with train-fitted geometry to evaluate output availability alongside surfaced-label error.
  • Evaluation findings: Across four within-dataset evaluations, evidence for geometry-augmented routing was favorable but not uniform.
Loading 2608.15951v1…