Source-linked AI summary
Generalizable machine learning for stress monitoring from wearable devices: A systematic literature review
Gideon Vos, Kelly Trinh, Zoltan Sarnyai, Mostafa Rahimi Azghadi
TL;DR
Stress varies across people, while large labeled wearable datasets for generic monitoring remain limited. This review synthesizes public stress datasets and associated machine-learning studies, assessing quality, validation, and generalization. It finds that study protocols and datasets vary widely, with shortcomings in labeling, statistical power, biomarker validity, and generalization.
Problem
Stress has subjective and biological components that vary between people, and large labeled wearable datasets for generic stress measurement remain limited.
Method
The review synthesizes published studies and public stress-biomarker datasets, using the IJMEDI checklist to assess study quality and model generalization.
Results
Only one reviewed machine-learning study was high quality, while the remaining studies were mostly medium quality, with one low-quality study.
Takeaways & Limitations
Reliable real-world stress monitoring requires further study of model generalization as newer and more substantial datasets become available.
Takeaways & Limitations
Public stress datasets commonly contain 25 or fewer subjects, limiting the evidence base for generalizable models.
Abstract
from arXiv · showhide
Introduction. The stress response has both subjective, psychological and objectively measurable, biological components. Both of them can be expressed differently from person to person, complicating the development of a generic stress measurement model. This is further compounded by the lack of large, labeled datasets that can be utilized to build machine learning models for accurately detecting periods and levels of stress. The aim of this review is to provide an overview of the current state of stress detection and monitoring using wearable devices, and where applicable, machine learning techniques utilized. Methods. This study reviewed published works contributing and/or using datasets designed for detecting stress and their associated machine learning methods, with a systematic review and meta-analysis of those that utilized wearable sensor data as stress biomarkers. The electronic databases of Google Scholar, Crossref, DOAJ and PubMed were searched for relevant articles and a total of 24 articles were identified and included in the final analysis. The reviewed works were synthesized into three categories of publicly available stress datasets, machine learning, and future research directions. Results. A wide variety of study-specific test and measurement protocols were noted in the literature. A number of public datasets were identified that are labeled for stress detection. In addition, we discuss that previous works show shortcomings in areas such as their labeling protocols, lack of statistical power, validity of stress biomarkers, and generalization ability. Conclusion. Generalization of existing machine learning models still require further study, and research in this area will continue to provide improvements as newer and more substantial datasets become available for study.
Generalizable Machine Learning for Stress Monitoring from Wearable Devices: A Systematic Literature Review
This review examines wearable-device stress detection with machine learning, emphasizing model generalization and reproducibility on unseen data. It concludes that most studies rely on small, singular datasets and require more varied evidence.
- The review analyzes machine-learning stress detection from wearable devices using the IJMEDI checklist and emphasizes generalization and reproducibility on unseen data.It synthesizes the current literature rather than presenting a new predictive model.
- Most stress-related machine-learning studies use small, singular datasets and lack generalization, motivating larger studies with substantially more varied datasets.
1. Introduction
Stress is a variable psychological and physiological response for which no universal evaluation standard exists. The review therefore focuses on wearable biomarkers, public datasets, machine-learning validation, and generalization to unseen data.
- Stress involves psychological and physiological responses, but a universally recognized standard for stress evaluation remains outstanding.Within the reviewed studies, stress is treated as a binary prediction condition.
- Wearable devices can measure biomarkers including HRV, EDA, and HR that may correlate with elevated stress.
- Earlier reviews did not address statistical power, labeling protocols, or whether models generalize across public datasets and conditions.
- The review investigates wearable-device stress measurement, public datasets, machine-learning approaches, accuracy, generalization, and limitations.
2. Methods
The review formulates questions about machine-learning techniques, reported accuracy, validation, reproducibility, and generalization. It searches published literature and assesses included studies with the IJMEDI checklist.
- 2.1. Research questions: The review asks which algorithms use public stress-biomarker datasets and whether reported findings are validated, reproducible, and promising for generalization.
- 2.2. Literature search: Searches of Google Scholar, Crossref, DOAJ, and PubMed identified 973 papers, with 16 duplicates removed before subsequent screening.
- 2.2. Literature search: The review excluded papers lacking detailed machine-learning, feature-engineering, or validation methods and ultimately selected 33 papers for systematic review.
- 2.3. Assessment of the quality of the studies: Two reviewers independently assessed study quality using the IJMEDI checklist covering data understanding, preparation, modeling, validation, and deployment.The checklist contains 30 questions across six dimensions and classifies scores from low to high quality.
3.1. Wearable Devices for Stress Measurement
The review focuses on stand-alone wearable devices and finds that studies predominantly use medical-grade or publicly available sensor data, especially from Empatica E4 devices. Most reviewed participants were healthy and health-screened, while single-biomarker devices were excluded.
- 3.1. Wearable Devices for Stress Measurement: The literature includes medical-grade and consumer-oriented wearables, but consumer devices generally do not provide raw biomarker recordings for scientific study.
- 3.1. Wearable Devices for Stress Measurement: The review limits wearable devices to stand-alone monitoring on the wrist, finger, or arm without a harness or paired secondary device.
- 3.1. Wearable Devices for Stress Measurement: Reviewed studies commonly use EDA and HR or HRV, while devices reporting stress from only a single biomarker were excluded.The review notes that biomarker validity remains an open research question.
- 3.1. Wearable Devices for Stress Measurement: The predominant wearable device in the reviewed publicly available datasets was the Empatica E4.
- 3.1. Wearable Devices for Stress Measurement: The vast majority of current studies used data from predominantly healthy subjects screened for health conditions before inclusion.
3.2. Wearable Device Datasets for Stress Measurement
Reviewed public wearable-device stress datasets use varied labeling and study protocols, but generally rely on small samples and laboratory or work-related stressors.
- Datasets: Public datasets predominantly contain EDA and HR biomarkers, with most including 25 or fewer subjects and Stress-Predict being largest at 35 subjects.AffectiveROAD and Toadstool each include 10 subjects; nearly all recorded sessions exceed 60 minutes except Toadstool.
- Labeling: Dataset labels were assigned either to predefined stressed/non-stressed intervals or to stress scores reported by participants or observers.The reviewed protocols therefore combine experimentally imposed conditions with subjective or observational assessments.
- Study protocols: Reviewed studies applied cognitive or work-related tasks under pressure, while only three collected stress biomarkers during normal life conditions.Subjects were screened for known health conditions in virtually all studies.
- Generalization: Laboratory-trained models outperformed models trained on normal-daily-life data when both predicted stress in daily-life conditions.This comparison highlights the importance of training context for models intended for real-world patient data.
3.3. Machine Learning Algorithms and Techniques for Stress Measurement using Wearable data
The reviewed stress-detection pipelines span sensor alignment, class balancing, feature engineering, algorithm selection, validation, and performance analysis. Reported performance is often high on labeled stress/no-stress periods, but class imbalance, information loss, and limited study quality constrain interpretation and generalization.
- Pre-processing: Class balancing was common because stress samples were generally less frequent, but neither up-sampling nor down-sampling substantially improved predictive power and discarding observations could lose biomarker information.Balancing can also reduce reproducibility and generalizability when new data contain outliers or different class distributions.
- Feature Engineering: Fourteen studies extracted sliding-window summary statistics, using windows from 0.25 seconds to 20 minutes, while feature-selection studies highlighted HR, respiratory rate, and HRV features.Window lengths and selected biomarkers varied across studies and protocols.
- Model Training and Validation: At least 64.5% test accuracy was achieved by the best model in each experiment, while several binary-classification studies reported over 90% using marked stress/no-stress periods.These accuracies measure correlation between wearable biomarkers and labels at the same time point, rather than necessarily forecasting future stress.
- Performance Analysis: Evaluation used varied metrics and validation procedures, including accuracy, F-score, Kappa, AUC, MAE, and Leave One Subject Out cross-validation.The review also reports that most included studies were medium quality, with an average quality score of 25.7 and biases in problem understanding, data understanding, and modeling.
4. Discussion
The review identifies recurring weaknesses in wearable stress-monitoring studies, including uncertain biomarker validity, labeling problems, limited statistical power, and weak generalization. These issues constrain confidence in reported accuracy and motivate larger, more rigorously validated datasets and models.
- 4. Discussion: Wearable stress models face four requirements: valid and varied biomarkers, accurate labels, sufficient statistical power, and generalization to unseen data.The review frames these requirements as necessary for robust stress detection and monitoring.
- 4. Discussion: Study quality was generally medium, with low scores in data preparation, deployment, and high-priority validation items, while quality did not notably improve over time.Only one study was rated high quality, and the review identified limited progress in modeling and little attention to real-world deployment, sustainability, bias, and ethics.
- 4.1. Validity of sensor biomarkers: Physiological signal validity remains uncertain because no standardized protocol exists, while motion artifacts can compromise wearable measurements during dynamic conditions.Empatica E4 data quality was reported as particularly vulnerable for EDA and IBI measurements under movement.
- 4.2. Labeling protocol: Labeling is also problematic: periodically labeled datasets produced higher detection accuracy than datasets using self-reported or observer-scored stress, potentially reflecting false-negative questionnaire reports.The review notes that stress is not inherently binary and that biomarker thresholds for low, moderate, and high stress were not established in the reviewed studies.
- 4.4. Lack of Generalization: Generalization evidence is scarce: nearly all studies used fewer than 30 subjects, and only two tested models on a completely unseen dataset.Figure 6 found no obvious correlation between subject count and reported accuracy; WESAD and SWELL corresponded to 45% and 70% power, respectively, under the stated assumptions.
5. Conclusion
This review synthesizes wearable-device stress prediction literature, including datasets, machine learning techniques, limitations, and generalization to unseen data. It identifies challenges and opportunities intended to advance stress detection and management technology.
- Its central objective is developing robust, accurate models that generalize well to new, unseen data.
- The review synthesizes wearable-device stress prediction studies, publicly available biomarker datasets, machine learning techniques, limitations, and generalization to unseen data.
- The review summarizes challenges and opportunities intended to advance machine learning for wearable-device stress detection and management technology.