Source-linked AI summary
Multisensor Measurement of Train Driver Mental Fatigue: From Simulation to Reality
Esther Bosch, Rebecca Kruschka, David Schackmann, Stephanie Hoyer, Wolfgang Kilian, Stefan Schwanitz, Anneke Hamann
TL;DR
Higher automation makes vigilance and situation awareness important monitoring concerns in railway operation. This study examined multisensor indicators across controlled and operational settings, finding the most consistent patterns for heart rate variability and breathing rate.
Problem
Vigilance and situation awareness can be degraded, motivating research into more direct and objective measures of operator state.
Method
The study examined subjective, physiological, and behavioral indicators of mental fatigue across controlled and operational rail settings using a multisensor approach.
Results
Heart rate variability and breathing rate showed the most consistent and theoretically coherent patterns across the controlled simulator and technically constrained real-world environment.
Takeaways & Limitations
Heart rate variability and breathing rate were the most consistent indicators across both examined environments.
Takeaways & Limitations
Traditional monitoring is limited when vigilance and situation awareness are degraded, motivating more direct objective measures.
Abstract
from arXiv · showhide
Increasing automation in rail transport shifts the train driver's role from active control to prolonged supervisory monitoring. This creates conditions for mental fatigue (MF) and reduced vigilance. Despite the safety relevance of this issue, evidence on the feasibility and robustness of physiological indicators of MF under operational rail conditions remains limited. Most prior work relies on simulators or lab studies. The present study investigated multiple subjective, physiological, and behavioral indicators of MF in professional train drivers across two complementary settings: a high-fidelity train simulator (n=14) and a real-world rail environment (n=6). To our knowledge, this is the first study to deploy a full multisensor battery under actual train operating conditions. In both settings, a standardized protocol was used comprising a baseline drive, a one-hour auditory n-back task as an MF induction procedure, and a second drive. Heart rate variability and breathing rate showed consistent and theoretically expected changes across both environments, suggesting reduced physiological arousal following the fatigue induction task. In contrast, EEG-based frontal theta power and parietal alpha and beta power, electrodermal activity, blink duration, and behavioral indicators did not show clear mental fatigue-related patterns. Real-world data collection revealed substantial technical challenges related to vibration, sensor connectivity, and concurrent high-frequency data acquisition. These findings suggest that autonomic indicators, particularly HRV and breathing rate, represent the most promising and ecologically robust measures for operational fatigue monitoring in train drivers. However, neurophysiological measures require further validation under realistic conditions before deployment in driver monitoring systems, and larger samples are needed to confirm these preliminary patterns.
1. Introduction
Increasing railway automation shifts drivers toward prolonged supervision, making mental fatigue relevant to safety-critical monitoring. The study evaluates which subjective, physiological, and behavioral indicators remain sensitive, feasible, and robust across simulator and real-world rail settings.
- Higher automation reduces active driving while requiring sustained monitoring, which can contribute to task-induced mental fatigue under monotonous conditions.
- Fatigue-related declines in sustained attention, situation awareness, action monitoring, and inhibitory control can impair responses in safety-critical monitoring tasks.
- Mental fatigue is difficult to measure objectively because it is subjective, multidimensional, and assessed indirectly through multimodal indicators.
- Railway research lacks a generally accepted standard for quantifying mental fatigue or selecting and combining suitable sensor systems.
- Operational measurement is constrained by lighting, vibration, sensor placement, comfort, privacy, and the need for non-intrusive robust systems.
- This study compares multiple fatigue indicators across a high-fidelity simulator and real-world rail operation to evaluate sensitivity, feasibility, and robustness.
2. Background
The background distinguishes mental fatigue from sleepiness and frames passive fatigue as a risk during prolonged monitoring. It reviews behavioral, ocular, EEG, autonomic, and multimodal approaches while emphasizing their limited operational validity.
- There is no universally accepted definition of mental fatigue, complicating comparisons across empirical studies.
- Mental fatigue develops gradually during sustained task engagement, whereas sleepiness and drowsiness arise from different mechanisms and can fluctuate more rapidly.
- Passive mental fatigue emerges during monotonous, low-demand monitoring and is associated with reduced engagement, mind wandering, slower hazard responses, and higher collision risk.
- 2.2. Traditional Mental Fatigue Monitoring and Its Limitations: Traditional acknowledgment systems can preserve correct responses despite degraded vigilance and situation awareness because drivers develop automated response patterns.
- The study addresses the need for evidence on which indicators remain interpretable, robust, and practically suitable in real-world rail environments.
3. Materials and Method
The study used a standardized fatigue-induction protocol in simulator and real-world rail settings, combining baseline and post-task drives with multisensor and behavioral measurements. Professional train drivers completed an auditory n-back task between drives.
- 3.1. Study design, Mental Fatigue Induction and Task Description: Sessions included a 15-minute baseline drive, a one-hour auditory n-back task, and a second 15-minute drive in both environments.
- 3.1. Study design, Mental Fatigue Induction and Task Description: The auditory n-back task presented three-digit numbers at four difficulty levels, requiring participants to enter target numbers on a keypad.
- 3.1. Study design, Mental Fatigue Induction and Task Description: Reaction time to matched orange trackside targets was measured in both environments as a behavioral indicator.
- 3.1. Study design, Mental Fatigue Induction and Task Description: The simulator approximated low-demand automated operation, whereas real-world driving occurred under GOA1 conditions with speed regulation and level-crossing management.
- The battery included subjective sleepiness ratings, EEG, eye tracking, HRV, breathing rate, electrodermal activity, and behavioral measures.
- The simulator sample comprised 14 professional drivers, while the real-world study used six professional drivers.
4. Results
Subjective sleepiness increased across drives in both environments, while autonomic measures showed consistent directional changes. EEG, electrodermal, blink-duration, and behavioral indicators largely lacked clear fatigue-related patterns, and none of the paired tests were significant.
- Behavioral and peripheral indicators: Electrodermal activity, blink duration, reaction time, and missed reminders did not provide clear fatigue-related evidence.Blink duration was nearly identical between drives, reaction times were faster in Drive 2, and missed-reminder distributions overlapped substantially.
- Subjective fatigue: KSS ratings increased from Drive 1 to Drive 2 in both environments, indicating greater subjective sleepiness after the n-back task.Simulator means increased from 3.21 to 5.50, while real-world means increased from 2.17 to 3.33.
- Neurophysiological indicators: Frontal theta, parietal alpha, and parietal beta power showed stability or decreases rather than the hypothesized fatigue-related increases.Frontal theta decreased in the real environment and remained nearly stable in the simulator; parietal beta decreased across drives.
- Physiological indicators: HRV increased from Drive 1 to Drive 2 in both environments, suggesting higher parasympathetic activity and lower physiological arousal.Simulator z-scored RMSSD changed from -0.35 to 0.32; the real-world sample changed from -0.31 to 0.32.
- Physiological indicators: Breathing rate decreased from Drive 1 to Drive 2 in both environments, with simulator means changing from 0.09 to -0.08.The real-world values were 0.13 in Drive 1 and -0.08 in Drive 2.
- Sensor comfort: Overall sensor comfort was rated favorably, although eye tracking was the least comfortable sensor.Overall sensor comfort averaged 3.85, while eye tracking averaged 3.55 on the five-point comfort scale.
5. Discussion
The discussion finds that subjective fatigue increased after the induction task in both settings, while baseline sleepiness and the size of the increase differed between environments. These differences warrant cautious interpretation because the real-world sample was smaller and the environments were not directly equivalent.
- RQ1: Fatigue induction: Simulator participants reported higher sleepiness before the session and a larger absolute increase than real-world participants.The discussion relates these differences to possible simulator-cab conditions, accumulated travel fatigue, or the more engaging real-world environment.
5.2. RQ2: Indicator Sensitivity
Indicator sensitivity was mixed: HRV and breathing rate followed hypothesized directions across environments, whereas EDA, EEG, blink duration, and behavioral measures did not clearly track the fatigue manipulation.
- Autonomic indicators: HRV and breathing rate matched hypothesized fatigue-related directions consistently across the simulator and real-world environments.RMSSD increased and breathing rate decreased from Drive 1 to Drive 2, consistent with reduced arousal.
- Autonomic indicators: The convergence of two autonomic indicators across both environments supports their potential robustness under differing operational conditions.The real-world pattern emerged despite a small and technically constrained sample.
- EDA: EDA did not show a consistent fatigue-related response: SCR peaks increased slightly in the real-world setting and remained essentially unchanged in the simulator.The real-world increase was opposite to the hypothesis and may reflect contextual demands or individual variability.
- EEG: EEG findings did not support the hypotheses, with frontal theta decreasing rather than increasing between drives.The authors interpret this non-replication as evidence that EEG fatigue markers may be context-sensitive under operational conditions.
- Blink duration: Blink duration remained nearly identical between drives, indicating insensitivity to the fatigue manipulation in this dataset.Task-relevant visual engagement may have suppressed fatigue-related blink lengthening, and only blink duration—not a composite index—was measured.
- Reaction time and dead man’s handle: Reaction times were faster in Drive 2 across both environments, while dead man’s handle misses showed no meaningful change in the simulator.The reaction-time pattern was attributed most likely to practice effects, whereas the automated reminder response may lack sensitivity.
5.3. RQ3: Ecological Robustness
The multisensor system was deployable in simulator and real-world rail settings, and HRV and breathing rate showed the same directional pattern despite substantial field constraints. However, small samples and environmental differences limit inferential and cross-setting conclusions.
- Ecological robustness: HRV and breathing rate showed the same directional pattern in the technically constrained real-world sample as in the controlled simulator.The authors describe this consistency as a meaningful form of robustness.
- Ecological robustness: Real-world data collection introduced substantial technical challenges that motivate purpose-built acquisition pipelines and systematic robustness testing.The discussion cautions against transferring laboratory sensor setups to operational rail environments without field validation.
- Study scope: Different sample sizes, participant composition, collection timing, and operational demands limit direct comparison between simulator and real-world environments.The study is therefore best understood as a feasibility investigation rather than a definitive environment comparison.
- Study scope: The real-world sample size was n=6, and most reported patterns should be treated as descriptive and exploratory.Larger samples and more controlled between-environment designs are needed before firm environment-specific conclusions can be drawn.
5.4. RQ4: Practical Feasibility
Drivers rated the sensors favorably overall, and the breathing belt combined comfort with the clearest fatigue-related sensitivity. Eye tracking showed the opposite pattern, being least comfortable and insensitive in blink-duration output.
- Most sensors received favorable usability ratings, with a median of 4 out of 5.
- The breathing belt was rated most comfortable and showed the clearest, most robust fatigue-related patterns.
- Eye tracking was least comfortable and its blink-duration output showed no fatigue sensitivity.
- These converging findings support prioritizing chest and abdomen worn autonomic sensors in future operational monitoring systems.
5.5. Returning to the Ironies of Automation
The behavioral findings illustrate the irony of automation: drivers maintained overt performance while subjective sleepiness increased. This supports a multisensor approach because autonomic measures do not depend on active response production.
- This pattern reflects automation’s reduction of active engagement while requiring reliable monitoring.
- Reaction times improved and dead man’s handle misses stayed flat across the session, even as subjective sleepiness rose.
- Drivers apparently maintained overt task performance through practiced, automatized responses while their underlying state shifted.
- Autonomic measures may offer a more direct route to detecting vigilance erosion than behavioral monitoring.
5.6. Adequacy of the Fatigue Induction
The auditory n-back task increased subjective fatigue, but its magnitude was modest, especially in the real-world environment. Its active cognitive demands may not represent the prolonged inactivity and underload relevant to operational monitoring.
- The procedure produced measurable physiological change within a 15-minute post-task drive.
- Subjective fatigue increased, but the magnitude was modest, particularly in the real-world environment.
- The n-back task induces active cognitive fatigue through sustained effortful processing.
- GoA2 monitoring fatigue instead arises from prolonged inactivity and underload, which may differ physiologically from n-back fatigue.
- A longer, more ecologically valid induction might target operational fatigue more directly and produce stronger physiological effects.
5.7. Practical Implications
HRV and breathing rate are the most promising current starting points for rail fatigue monitoring because they changed consistently across environments, were technically robust, and were comfortable for drivers. Other measures require further validation, while future studies need larger samples and field-ready infrastructure.
- HRV and breathing rate showed consistent directional changes across environments and were rated among the most comfortable sensors.
- EEG showed no consistent fatigue-related pattern and faces substantial practical challenges in field deployment.
- EEG, EDA, blink duration, and the tested behavioral measures were not sensitive to the manipulation under the studied conditions.
- Future work should use ecologically valid fatigue inductions, larger samples, and purpose-built acquisition infrastructure.
- Sensor comfort and wearability require systematic investigation because long-term driver acceptance is a prerequisite for operational monitoring.
6. Conclusion
The study found that heart rate variability and breathing rate showed the most consistent fatigue-related patterns across simulator and real-world rail settings. Other physiological and behavioral indicators were less clear, while field deployment exposed practical constraints requiring further validation and larger samples.
- The study examined multisensor mental-fatigue assessment in professional train drivers across simulator and real-world rail settings.The approach was intended to evaluate indicator feasibility, sensitivity, and robustness across complementary environments.
- Subjective sleepiness increased across the session in both settings, indicating that the fatigue induction procedure was effective at least at the self-report level.
- Heart rate variability and breathing rate showed the most consistent and theoretically coherent patterns across simulator and technically constrained real-world environments.Both sensors were also rated among the most comfortable by drivers.
- EEG, electrodermal activity, blink duration, and behavioral measures did not show clear fatigue-related changes.The failure to replicate expected frontal theta increases underscores the context-sensitivity of neurophysiological fatigue markers.
- Reaction times and dead man’s handle performance improved or stayed flat despite rising subjective fatigue, illustrating a gap between behavioral output and internal state.This pattern supports the relevance of autonomic monitoring for tracking driver state when overt task performance remains sustained.
- Real-world data collection revealed challenges involving vibration, sensor connectivity, and concurrent high-frequency data acquisition.Future work should use more ecologically valid fatigue induction, larger samples, and purpose-built acquisition infrastructure for operational rail environments.