Source-linked AI summary
Neural network an1alysis of sleep stages enables efficient diagnosis of narcolepsy
Jens B. Stephansen, Alexander N. Olesen, Mads Olsen, Aditya Ambati, Eileen B. Leary, Hyatt E. Moore, Oscar Carrillo, Ling Lin, Fang Han, Han Yan, Yun L. Sun, Yves Dauvilliers, Sabine Scholz, Lucie Barateau, Birgit Hogl, Ambra Stefani, Seung Chul Hong, Tae Won Kim, Fabio Pizza, Giuseppe Plazzi, Stefano Vandi, Elena Antelmi, Dimitri Perrin, Samuel T. Kuna, Paula K. Schweitzer, Clete Kushida, Paul E. Peppard, Helge B. D. Sorensen, Poul Jennum, Emmanuel Mignot
TL;DR
Manual sleep-stage scoring is time-consuming, subjective, and inconsistent, motivating an automated alternative for diagnosing Type-1 Narcolepsy. The paper trains neural networks on diverse sleep recordings, introduces hypnodensity-based staging, and develops a single-night biomarker that matches the clinical gold standard while supporting finer-resolution scoring and potential home use.
Problem
Manual polysomnography scoring is time-consuming, expensive, subjective, and inconsistent, with especially low agreement for N1 and N3 stages.
Method
Neural networks were trained and evaluated on several thousand sleep studies from 10 cohorts across 12 sleep centers, producing probabilistic hypnodensity graphs and a narcolepsy biomarker from sleep-stage overlaps.
Results
The ensemble achieved 87% accuracy and outperformed individual scorers, scored stages at 5-second resolution, and produced a T1N biomarker comparable to MSLT using a single sleep study.
Takeaways & Limitations
The method could reduce clinic time and cost, simplify T1N diagnosis, and enable diagnosis from home sleep recordings.
Takeaways & Limitations
A broader universal sleep-analysis model would require additional subjects with narcolepsy and other conditions in the training data.
Abstract
from arXiv · showhide
Analysis of sleep for the diagnosis of sleep disorders such as Type-1 Narcolepsy (T1N) currently requires visual inspection of polysomnography records by trained scoring technicians. Here, we used neural networks in approximately 3,000 normal and abnormal sleep recordings to automate sleep stage scoring, producing a hypnodensity graph - a probability distribution conveying more information than classical hypnograms. Accuracy of sleep stage scoring was validated in 70 subjects assessed by six scorers. The best model performed better than any individual scorer (87% versus consensus). It also reliably scores sleep down to 5 instead of 30 second scoring epochs. A T1N marker based on unusual sleep-stage overlaps achieved a specificity of 96% and a sensitivity of 91%, validated in independent datasets. Addition of HLA-DQB1*06:02 typing increased specificity to 99%. Our method can reduce time spent in sleep clinics and automates T1N diagnosis. It also opens the possibility of diagnosing T1N using home sleep studies.
INTRODUCTION
Sleep disorders are commonly evaluated from manually scored polysomnography, but this process is laborious, subjective, and inconsistent. The paper proposes neural-network scoring that preserves uncertainty through hypnodensity graphs and supports automated narcolepsy analysis.
- INTRODUCTION: Polysomnography records multiple physiological signals and assigns each epoch to wake, N1, N2, N3, or REM.The paper describes PSG as including EEG, EOG, EMG, ECG, breathing effort, oxygen saturation, and airflow.
- INTRODUCTION: Manual sleep staging is time consuming, expensive, subjective, inconsistent, and generally performed offline.Average inter-scorer reliability was 82.6%, with agreement as low as 63% for N1 and 67% for N3.
- INTRODUCTION: Deep learning was investigated as a fast, inexpensive, objective, and reproducible alternative to manual sleep-stage scoring.The motivation follows reported progress of neural networks in image labeling, speech understanding, translation, and healthcare applications.
- INTRODUCTION: The hypnodensity graph assigns membership functions to sleep stages instead of enforcing one label, conveying more information about sleep trends.This probabilistic representation is presented as possible through non-human scoring.
- INTRODUCTION: The study used six-scorer consensus data to evaluate automated scoring and address disagreements among human scorers.The inter-scorer reliability cohort contained 70 PSGs scored by six scorers across three US locations.
Optimizing machine learning performance for sleep staging
The study optimized neural-network sleep staging across cohorts, model architectures, signal encodings, and segment settings. An ensemble quantified uncertainty and matched or exceeded human-scoring benchmarks, while performance varied with dataset characteristics and model choices.
- Optimizing machine learning performance for sleep staging: Model accuracy varied across datasets, with lower performance in cohorts containing more fragmented sleep and higher performance where labels were more accurate.The worst performance occurred in KHC and SSC narcolepsy data, while IS-RC produced the best performance.
- Optimizing machine learning performance for sleep staging: Encoding and memory were the two most important factors increasing prediction accuracy, while segment length, complexity, and realizations were less important.Cross-correlation encoding benefited from higher complexity, whereas octave encoding worsened; memory helped octave models more than cross-correlation models.
- Optimizing machine learning performance for sleep staging: Hypnodensity outputs represented stage probabilities per epoch and showed greater uncertainty where scorer consensus was lowest.The ensemble came closest to the scoring consensus in the example comparison.
- Optimizing machine learning performance for sleep staging: The final scorer used an ensemble of cross-correlation models with varied parameters to reduce noise and quantify uncertainty.Sixteen models were trained, and mean and variance were calculated for each analyzed segment.
- Optimizing machine learning performance for sleep staging: Weighted performance increased overall accuracy from 87% to 94% by down-weighting segments with controversial scorer assessments.The largest discrepancies involved wake versus N1, N1 versus N2, and N2 versus N3.
Influences of sleep pathologies
The study evaluated how sleep pathologies affect automated staging and developed a narcolepsy biomarker from neural-network sleep-stage outputs. Performance was tested across cohorts, with independent testing showing strong sensitivity and specificity.
- Accuracy varied across cohorts and was analyzed against cohort, age, sex, insomnia, OSA, RLS, PLMI, and T1N.
- Five-second model segments were averaged to evaluate staging at 5-, 10-, 15-, and 30-second resolutions.
- Sleep-stage mixing and dissociation in T1N motivated using overlapping neural-network stage probabilities as diagnostic features.
- The biomarker used geometric-mean time-series features from 16 sleep-stage prediction models, supplemented with REM latency and sequencing features.
- 91% sensitivity and 96% specificity were achieved in testing data, while replication reached 93% sensitivity and 91% specificity.
- 90% sensitivity and 99% specificity were achieved after adding HLA typing in the test and replication evaluation.
DISCUSSION
The discussion presents neural-network hypnodensity scoring as an accurate, interpretable approach for automated sleep staging and T1N detection. It emphasizes clinical efficiency while retaining important boundaries around interpretation and MSLT replacement.
- Machine learning scored PSG sleep stages accurately across locations, recording environments, protocols, hardware, software configurations, and sleep disorders.
- Ensembling multiple algorithms increased robustness because individual models respond differently to aspects of each recording.
- 87% performance was estimated against consensus scoring from six technicians across 70 subjects, with the algorithm outperforming individual scorers.
- Sleep-stage disagreements were concentrated at wake–N1, N1–N2, and N2–N3 transitions, where stage boundaries are subjective.
- Hypnodensity distributions encode stage probabilities and yielded interpretable features for T1N diagnosis from a single PSG night.
- The biomarker achieved performance similar to MSLT while requiring one sleep study, and HLA typing raised specificity above 99% without loss of sensitivity.
- The algorithm does not replace MSLT for measuring daytime sleepiness through mean sleep latency across naps.
- A direct multitask classifier might improve predictions but would reduce feature interpretability and require more subjects with narcolepsy and other conditions.
METHODS
The models were trained, validated, and tested on several thousand sleep studies drawn from heterogeneous cohorts and multiple international sleep centers. One longitudinal cohort supplied most training data and additional validation studies.
- Data came from 10 cohorts recorded at 12 sleep centers across three continents.
- The Wisconsin cohort contributed 2,167 PSGs from 1,086 subjects for training and 286 randomly selected PSGs for validation testing.
Patient-based Stanford Sleep Cohort
The patient-based cohorts comprised clinical and high-pretest-probability samples from multiple sites, with datasets allocated across staging validation, biomarker training, and independent testing.
- The Stanford cohort contributed 894 diagnostic PSG recordings from independent patients with diverse sleep diagnoses.
- The Korean Hypersomnia Cohort included 160 patients with excessive daytime sleepiness and was used for testing sleep scoring and training the narcolepsy biomarker.
- The Austrian cohort contained 118 PSGs from 86 high-pretest-probability patients, including 42 clear T1N cases with cataplexy.
- The inter-scorer reliability cohort comprised 70 PSGs scored by six scorers across three United States locations.
- The sodium-oxybate trial sample provided seven baseline PSGs from patients with clear and frequent cataplexy for biomarker training.
- The Italian cohort included 70 T1N patients and 77 other patients with hypersomnia-related conditions, with subjects divided between training and testing.
- The Danish cohort contained 79 PSGs from controls and patients classified using PSG, MSLT, and CSF hypocretin-1 measures.
Patient-based French Hypersomnia Cohort
The cohorts comprise French PSG recordings from T1N, other narcolepsy, and control participants, alongside an inter-scorer reliability dataset. Sleep scoring uses technician-assigned 30-second epochs, while the model produces probabilistic stage outputs that preserve within-epoch variation.
- Patient-based French Hypersomnia Cohort: 122 PSGs in the French Hypersomnia Cohort included 63 T1N subjects and 22 narcolepsy type 2 subjects.
- Patient-based French Hypersomnia Cohort: 199 PSGs in the Czech National Cohort included 67 T1N subjects with clear-cut cataplexy and 132 randomly selected population controls.
- Patient-based French Hypersomnia Cohort: The AASM ISR dataset contains one control study of 150 30-second epochs scored by 5234±14 experienced sleep technologists for quality control.
- Data labels and scoring: Technicians assign each 30-second PSG epoch a single stage label using established AASM scoring rules.The label is based on the stage judged to occupy the majority of the epoch.
- Data labels and scoring: The hypnodensity graph represents a probability distribution over possible sleep stages for each epoch rather than enforcing one label.For comparison with the gold standard, the model output is converted to a discrete estimated label.
Data selection and pre-processing
The preprocessing pipeline selects relevant PSG channels and transforms noisy biophysical signals into representations suited to convolutional neural networks. Octave encoding preserves multiscale frequency information, while cross-correlation encoding emphasizes periodic structure and attenuates noise.
- Channel selection: EEG, chin EMG, and bilateral EOG channels were selected from full-night PSG recordings for sleep scoring.The study notes that some PSG channels are unnecessary and that poor electrode connections can make recordings noisy.
- Signal representations: Octave encoding repeatedly low-pass filters each channel to produce five new channels while retaining the original signal information.High-frequency information can be recovered by subtracting lower-frequency channels.
- Signal representations: Scaling each channel to its 95th percentile and applying a log-modulus transform attenuates very large values from noisy regions.The scaling reference is derived from overlapping 90-minute segments rather than the entire recording.
- Signal representations: Octave decomposition applied across channels produces 25 new channels in total.
- Signal representations: Cross-correlation encoding reveals underlying periodicities while attenuating uncorrelated noise in PSG signals.For a signal compared with an extended version of itself, zero lag represents segment power.
- Signal representations: Cross-correlation represents frequency content as oscillation patterns that CNNs can detect across an input, unlike localized spectrogram features.The length and size of the correlation function reflect expected frequency content and quasi-stationarity.
Architectures of applied CNN models
The applied CNN systems use separate modality-specific subnetworks for EEG, EOG, and EMG, followed by layers that combine their outputs. The study varies model complexity, memory, and ensembling while using regularization during training.
- Model architecture: Three separate subnetworks process EEG, EOG, and EMG before fully connected layers combine their inputs into a softmax output.Memory-based models replace fully connected hidden units with LSTM cells and recurrent connections between successive segments.
- Model architecture: Two model sizes were evaluated to quantify the effect of increasing architectural complexity.More complex models have more parameters and are more likely to over-fit, although regularization can reduce this risk.
- Training: The training objective uses cross-entropy with L2 regularization, with the weight-decay parameter set to 0.00001.Parameters were initialized from N(0,0.01) and optimized with stochastic gradient descent with momentum.
- Training: Momentum was set to 0.9 and the learning rate initially to 0.005, followed by exponential learning-rate decay.
- Regularization: Batch normalization, weight decay, early stopping, and LSTM dropout set at 0.5 were used to reduce over-fitting.Validation used 10% of the training data and occurred after every 50th training batch.
- Ensembling: Ensembling formed sleep-stage predictions from multiple model predictions to target higher predictive performance than a single model.
Performance comparisons of generated CNN models
The study evaluates model factors and constructs T1N features from hypnodensity-derived sleep-stage patterns. Features summarize stage combinations, temporal behavior, and abnormal transitions across recordings.
- Model comparison: A 2^5 factorial experiment produced 32 models varying encoding, segment length, complexity, memory, and model realizations.Models were compared on a per-epoch basis.
- T1N feature construction: Hypnodensity-based T1N features use geometric means for every permutation of the five sleep stages: wake, REM, N1, N2, and N3.There are 31 nonempty stage combinations, with features derived from their temporal behavior.
- T1N feature construction: For each stage-combination size, fifteen features based on the mean, derivative, entropy, and cumulative sum were extracted.
- T1N feature construction: Transition features calculate the geometric mean between successive hypnodensity peaks and combine transitions of the same type.Wake and N1 peaks were merged, and peaks below 10 were excluded to reduce spurious detections.
Gaussian process models for narcolepsy diagnosis
The study uses Gaussian process classifiers to make probabilistic, nonlinear narcolepsy predictions while estimating uncertainty and incorporating HLA-DQB1*06:02 status.
- Gaussian process classification: Gaussian process classifiers provide nonlinear decision boundaries and uncertainty estimates for combining sleep-derived features.The models are non-parametric probabilistic classifiers using kernels.
- Feature selection: 38 features were selected after recursive feature elimination for use in the Gaussian process classifier.Feature selection was used to reduce overfitting and improve interpretability.
- HLA integration: HLA-DQB1*06:02 was encoded as a binary predictor, producing negative narcolepsy predictions when the test result was negative.The marker is positive in 97% of patients meeting biochemical or clinical disease definitions.
High pretest probability sample
The study evaluates automatic scoring and narcolepsy detection across clinical datasets, using heterogeneous cohorts, model pipelines, and diagnostic performance analyses.
- High pretest probability sample: The high-pretest analysis evaluates whether the detector distinguishes T1N from other causes of unexplained sleepiness.The comparison includes patients assessed through MSLT-based clinical workups.
- Evaluation: ROC analyses compare sensitivity-specificity tradeoffs across training, testing, replication, and high-pretest samples, with and without HLA.The figure marks cutoff thresholds for both model variants.
- Study design: The study pipeline preprocesses PSG signals, trains neural networks, extracts hypnodensity features, and evaluates a Gaussian process classifier.Training and testing use separated data splits, with testing features extracted using the same procedure.
- Model inputs: Neural-network inputs use processed EEG, EOG, and EMG signals represented through octave or cross-correlation encodings.The supplied figure legend identifies the two data formats but does not provide the complete network architecture.
TABLES
The tables document scorer performance, model comparisons, ensemble confusion patterns, and selected sleep-derived features used in narcolepsy classification.
- Table 1: Table 1 reports individual and overall scorer accuracy and Cohen’s kappa using biased and unbiased leave-one-out consensus standards.It also reports paired t-statistics and p-values comparing unbiased scorer predictions with model predictions.
- Table 2: Table 2 compares best-model performance across datasets against the six-scorer consensus on an epoch-by-epoch basis.The supplied table description does not include the numerical values.
- Table 3: Table 3 compares ensemble estimates with unweighted and scorer-agreement-weighted consensus targets across 53009 and 36032 epochs.Weighting emphasizes epochs with greater scorer agreement.
- Table 4: Frequently selected narcolepsy features include sleep-stage timing, REM-distribution entropy, maximum wakefulness probability, and weighted overlap measures.These features represent sleep-stage dissociation, REM consolidation, wakefulness, and accumulated probabilistic stage overlap.