Source-linked AI summary

Learning transferable human physiology from two million hours of sleep with SleepFM-2

Rahul Thapa, Christopher Sun, William Theodor Lehn-Schioler, Sophia Claire Kivelson, Umaer Hanif, Hyatt Moore, Harrison G. Zhang, Hafsa Ahmed, Marcus Dige, Niels R. Lorenzen, Elisabeth Roxane M. Heremans, Adrien Specht, Ulysse Gimenez, Robin Guillard, Andreas Brink-Kjaer, James Zou, Emmanuel Mignot

arXiv:2609.06849v1cs.AI

TL;DR

SleepFM-2 asks whether multimodal physiology captured during clinical sleep can yield a transferable representation of health beyond conventional PSG summaries and laboratory settings. The paper pretrains a frozen encoder on large-scale multimodal PSG and evaluates it across disease prediction, sleep assessment, subjective endpoints, and wearable sensors. Across these settings, the representation transfers across diseases, clinical tasks, sensing modalities, and aspects of subjective sleep, while its clinical utility and subgroup consistency remain unresolved.

  • Problem

    It remains unclear whether PSG foundation-model representations capture subjective sleep, shared disease-related physiology, and signals that generalize beyond pretrained health systems and sensors.

  • Method

    SleepFM-2 uses a frozen multimodal PSG encoder pretrained on more than two million hours of data and evaluates fixed embeddings across clinical tasks, disease outcomes, subjective endpoints, and wearable sensing.

  • Results

    SleepFM-2 performs within the range of expert scorers, predicts disease across two independent health systems, outperforms interpretable same-recording features, and transfers to lower-burden sensors including wrist accelerometry.

  • Takeaways & Limitations

    Multimodal clinical sleep physiology can provide a transferable representation of human physiology across diseases, clinical tasks, sensing modalities, and subjective experience.

  • Takeaways & Limitations

    Disease analyses are retrospective and observational, and improved discrimination alone does not establish clinical utility; subgroup performance was also not comprehensively assessed.

Abstract

from arXiv · show

Sleep provides a nightly window into health by capturing coordinated activity across the brain, heart, muscles and respiratory system. We introduce SleepFM-2, a sleep foundation model developed and evaluated on 282,511 polysomnography recordings from 26 cohorts, including 235,865 used for pretraining. These data span more than two million hours of multimodal physiology. Compared with SleepFM, SleepFM-2 improves disease prediction and sleep scoring, supports arousal, limb movement and respiratory event detection, and transfers to wearable sensing and subjective sleep phenotypes. A model combining its PSG representation with age, sex and BMI met a prespecified discrimination and significance criterion for 215 subsequently recorded EHR phenotypes in two held-out cohorts, including one health system unseen during pretraining. For 155 phenotypes, the PSG representation added reproducible information beyond demographics. SleepFM-2 also outperformed a 480-feature baseline derived from the same recordings. Its disease scores revealed a reproducible principal component associated with reduced sigma-band spatial coupling and increased hypnodensity entropy. The frozen encoder performed within the observed range of expert scorers for sleep events and transferred to wakeful EEG, headband and in-ear EEG, wrist PPG and wrist accelerometry. It improved sleep staging across six accelerometry cohorts and achieved disease-prediction performance in UK Biobank similar to models pretrained directly on accelerometry. Finally, SleepFM-2 captured aspects of subjective sleep not recovered by conventional PSG summaries, particularly reports of the recorded night. These results show that multimodal sleep physiology can provide a transferable representation of human health across diseases, clinical tasks, sensors and subjective experience.

Introduction

SleepFM-2 addresses whether multimodal physiology from a single sleep recording can support broader clinical, subjective, and wearable-sensing tasks. It uses a frozen representation learned from multimodal PSG to evaluate disease prediction, fine-grained sleep scoring, subjective endpoints, and transfer beyond the sleep laboratory.

  • Open questions: Existing PSG foundation models leave open whether learned representations capture subjective sleep, common disease-related physiology, and signals that generalize to unseen health systems.These questions arise because medical models may learn site-specific structure that does not transfer, while subjective self-assessment and shared disease dimensions remain uncertain.
  • Clinical scope: SleepFM-2 extends PSG assessment beyond staging to cortical arousals, limb movements, and respiratory events by operating at finer temporal resolution.These events determine clinically used indices and occur at scales too fine for SleepFM’s five-second tokenization.
  • Evaluation scope: The frozen encoder is evaluated on incident disease prediction, sleep scoring, subjective sleep and symptom endpoints, and transfer to signals outside the sleep laboratory.The evaluation includes wakeful EEG, headband and in-ear EEG, wrist PPG, and wrist accelerometry, including a sensor with no shared PSG channel.
  • Contribution: SleepFM-2 learns a multimodal physiology representation that transfers across clinical outcomes, tasks, and sensing modalities.This includes lower-burden wearable sensors and accelerometry reached through an engineered bridge using respiratory and cardiac surrogates.
  • Motivation and approach: SleepFM-2 is pretrained on more than two million hours of multimodal PSG and evaluated across disease, sleep-scoring, subjective, and wearable-sensing tasks.The model is trained on 235,865 recordings from 24 PSG cohorts and supports transfer to headband and in-ear EEG, wrist PPG, and actigraphy.

SleepFM-2 Overview

SleepFM-2 is a next-generation multimodal PSG foundation model evaluated on held-out cohorts and sensors. Its one-second, per-channel design supports finer-grained event scoring while a single frozen encoder serves multiple downstream tasks.

  • Dataset and evaluation: SleepFM-2 is pretrained and evaluated on more than two million hours of PSG from 26 independent cohorts, with two cohorts held out of pretraining.The held-out datasets evaluate transfer to an unseen health system, population, and sensor family.
  • Downstream tasks: A single frozen encoder supports sleep stages, cortical arousals, limb movements, respiratory events, phenome-wide disease prediction, and subjective sleep characterization.These downstream tasks are scored at one-second resolution for the event-detection tasks.
  • Architecture: One-second tokenization enables modeling shorter arousals, limb movements, and respiratory dynamics that five-second tokens cannot represent adequately.The finer resolution expands coverage beyond staging to three additional clinically relevant event tasks.
  • Architecture: Per-(channel, patch) tokens preserve anatomical and channel-specific information while providing embeddings at token, pooled-per-second, and modality-summary resolutions.This design allows representations to retain information before immediate cross-channel pooling.
  • Ablations: The encoder shrinks from 4.83M to 2.57M parameters while performance improves across the three downstream tasks in the fixed-corpus ablation chain.The comparison holds the pretraining corpus and token size fixed so that only the model varies.

Disease prediction from a single night of sleep

SleepFM-2 predicts incident disease across diverse organ systems from a single night of PSG, with reproducible gains beyond demographics and interpretable feature baselines in held-out cohorts. Its representation also improves sleep staging and performs within the range of expert scorers for finer-grained sleep events.

  • Phenome-wide prediction: 215 phenotypes met the prespecified C-index and significance criterion in both held-out cohorts, while 155 added reproducible PSG information beyond demographics.The stricter criterion required a Cox likelihood-ratio test for incremental PSG signal.
  • Phenome-wide prediction: SleepFM-2 predicted cardiovascular, respiratory, renal, metabolic and neurological diseases, including heart failure, atrial fibrillation, chronic kidney disease and diabetic complications.These diseases span systems directly and indirectly represented in PSG physiology.
  • Phenome-wide prediction: SleepFM-2 reached C-index values of 0.918 and 0.886 for Parkinson’s disease across SSC and HSP, respectively, with many other conditions showing reproducible discrimination.Among conditions meeting additional performance and significance requirements, 151 exceeded demographics by at least 0.03 in both cohorts.
  • Baseline comparisons: Against the strongest interpretable baseline, the 480-feature bank with demographics, SleepFM-2 reached C-index values of 0.723 and 0.711 versus 0.715 and 0.689.The embedding was ahead for 67% of conditions in SSC and 86% in HSP.
  • Baseline comparisons: Mean disease C-index increased from 0.709 to 0.718 in SSC and from 0.696 to 0.729 in HSP when comparing SleepFM-1 with SleepFM-2.The larger HSP gain occurred in a health system neither model had seen during training.
  • Clinical sleep scoring: SleepFM-2 improved sleep-staging macro F1 across cohorts and exceeded the average expert scorer on cortical arousal, limb movement and respiratory event detection.Mean F1 was 0.601 versus 0.546, 0.604 versus 0.530 and 0.462 versus 0.404, respectively; each task remained within the expert range.
  • Shared physiological signal: Disease risk scores formed a shared low-dimensional axis: the leading component explained 57.3% of variance in SSC and 58.2% in HSP.The axis was associated with reduced sigma-band spatial coupling and more ambiguous hypnodensity patterns.

SleepFM-2 transfers to wakeful EEG, wearable EEG and PPG, and wrist accelerometry

SleepFM-2 transfers from multimodal PSG to wakeful EEG, wearable EEG and PPG, and wrist accelerometry. Its frozen encoder improves sleep staging across wearable settings and matches directly pretrained accelerometry models for disease prediction.

  • Wakeful EEG: 0.163 mean AUROC gain across 23 wakeful EEG datasets was achieved over the baseline, with the largest gain in clinical and pathological detection.The mean gain was 0.205 across 12 clinical and pathological datasets, 0.148 across seven cognitive, affective and phenotyping datasets, and 0.062 across four motor-imagery datasets.
  • Wearable EEG and PPG: Macro F1 rose from 0.514 to 0.721 on BOAS headband EEG, from 0.463 to 0.645 on Wearanize+ headband EEG, and from 0.239 to 0.746 on EESM19 ear EEG.On Wearanize+ wrist PPG, F1 rose from 0.175 to 0.531.
  • Wrist accelerometry: 0.708–0.852 five-class AUROCs across six accelerometry cohorts exceeded E2E-Supervised results of 0.590–0.690 on every cohort.The pretrained encoder’s improvement over baseline ranged from 0.106 to 0.221.
  • Wrist accelerometry: 0.687 mean C-index across 388 UK Biobank outcomes matched the 0.688 achieved by accelerometry foundation models on the matched test cohort.Both models used age, sex and BMI; the matched test cohort included 5,192 participants.
  • Wrist accelerometry: 0.6913 mean C-index for an untuned ensemble of both representations exceeded 0.6882 for accelerometry foundation models by +0.0028.The paired difference was +0.0028 with a 95% CI of +0.0003 to +0.0053, although the improvement was small.

SleepFM-2 predicts poor subjective sleep and symptom endpoints

SleepFM-2 extracts information about subjective sleep, symptoms and discrepancies from raw PSG beyond conventional summaries and reference features. Its strongest subjective-sleep performance concerns the recorded night rather than habitual sleep.

  • Recorded-night sleep: SleepFM-2 was the best method for all six Post-Sleep Questionnaire items, improving AUROC by +0.011 to +0.034 over the reference methods.Mean AUROC was 0.750 versus 0.725, and the improvement survived Benjamini–Hochberg correction on four of six endpoints.
  • Symptoms and comorbidities: 0.759 versus 0.750 mean AUROC across four medical-history comorbidities made SleepFM-2 the best-performing method for all four endpoints.It significantly improved depression by +0.015.
  • Symptoms and comorbidities: 0.678 versus 0.657 mean AUROC across four symptom endpoints made SleepFM-2 the best-performing method for insomnia, fatigue, parasomnia and somatic symptoms.It significantly outperformed Combined for insomnia, parasomnia and somatic symptoms, while performing comparably to Combined for fatigue.
  • Subjective–objective discrepancy: 0.646 AUROC for |∆TST| and 0.748 for |∆SOL| exceeded Combined by +0.024 and +0.020, respectively.These results indicate that raw physiological signals contain information about subjective–objective discrepancy beyond conventional PSG summaries and individual microstructural features.
  • Composite sleep endpoints: 0.774 versus 0.745 AUROC on the Last-night composite produced a paired improvement of +0.028 over Combined.For the Habitual composite, performance was essentially indistinguishable: 0.680 versus 0.677, with Δ = +0.002 and p = 0.82.

An interactive web interface for SleepFM-2

SleepFM-Interface provides an interactive, no-programming pathway for inspecting SleepFM-2 outputs from uploaded PSG recordings. It is presented as research workflow software rather than a validated clinical readout.

  • Interface workflow: SleepFM-Interface maps uploaded PSG channels to the encoder vocabulary, produces frozen-encoder embeddings, and returns task-specific outputs through an interactive webpage.Users can inspect or override processing steps, request explanations, and draft reports.
  • Interface workflow: The interface includes summary, recording and details views, with measurements underlying the summary available for verification.The screenshots demonstrate the workflow on a public, de-identified CAP Sleep Database recording not included in pretraining.
  • Scope: The tool is research software and is not validated for clinical use, so prospective clinical evaluation remains future work.The demonstration is explicitly not a clinical readout.

Discussion

SleepFM-2 extends PSG-based sleep modeling from conventional scoring toward transferable representations of disease risk, subjective sleep, and wearable physiology. Its benefits are broad but bounded by retrospective clinical evaluation, PSG-selected cohorts, and the limits of single-night recordings for longitudinal sleep constructs.

  • Clinical interpretation: SleepFM-2 performed within expert-scorer ranges on multiple sleep-scoring tasks and extended automated interpretation to arousals, limb movements, and respiratory events.Its finer temporal resolution and expanded event coverage address tasks beyond sleep staging.
  • Disease prediction: Phenome-wide predictions exceeded interpretable features derived from the same recordings, with the advantage clearest in the health system unseen during pretraining.The margin was narrow in SSC and roughly three times larger in the unseen health system.
  • Shared physiology: A common physiological-risk component explained approximately 57 to 58% of variance in disease-specific predictions and was associated with sigma-band spatial coupling and hypnodensity entropy.The authors interpret this as consistent with a latent dimension of physiological aging or frailty, while calling for validation of its clinical value.
  • Subjective sleep: Subjective-sleep performance was strongest for the recorded night, whereas habitual-sleep performance was less consistent and did not replace longitudinal assessment.SleepFM-2 outperformed the combined baseline on all six recorded-night questionnaire items but was significantly better on only two of six matched sleep-history items.
  • Wearable transfer: A frozen PSG encoder transferred to headband and in-ear EEG, wrist PPG, and wrist accelerometry, including disease prediction despite no shared PSG channel.Its wrist-accelerometry disease-prediction performance was comparable to models pretrained directly on accelerometry in UK Biobank.
  • Limitations: Disease analyses were retrospective and observational, and the primary cohorts comprised people referred for laboratory PSG, limiting conclusions about clinical utility and broader populations.Prospective studies are needed to evaluate downstream decisions, calibration, subgroup performance, and people who would not typically undergo PSG.

Methods

SleepFM-2 jointly encodes four PSG modality groups using a shared encoder trained with contrastive and masked-reconstruction objectives. Its architecture preserves channel and temporal structure while exposing embeddings at resolutions matched to downstream sleep and event-detection tasks.

  • Architecture: SleepFM-2 jointly encodes brain and eye activity, cardiac, muscle, and respiratory PSG streams into a shared embedding space.The model uses a leave-one-out contrastive objective together with masked autoencoding of individual channel-patch tokens.
  • Input representation: PSG inputs are organized as four modality streams with fixed channel-set alignment, preserving anatomical channel identity and tracking missing channels through attention and pooling.The input tensors have shape (B, C_m, T), with PAD positions explicitly propagated.
  • Tokenization: A shared strided 1D CNN tokenizer converts each channel into non-overlapping temporal patch tokens, while learned channel-region embeddings supply anatomical identity.Channel identity is not encoded implicitly in tokenizer parameters and is added separately.
  • Encoder: Stage 1 applies transformer blocks jointly across channel and time tokens, operating on visible non-PAD tokens during pretraining and the full grid at inference.This design combines cross-channel and cross-time context while using masking to reduce pretraining sequence length.
  • Embedding hierarchy: The contrastive branch pools channel information at each patch, refines temporal context, and produces per-modality summary embeddings alongside finer token and patch representations.Downstream tasks select patch-level, summary-level, or token-level embeddings according to target granularity.
  • Pretraining objectives: The masked-autoencoding branch reconstructs masked non-PAD signal patches, and the total pretraining loss is L = λ_CL · L_CL + λ_MAE · L_MAE with both weights set to 1.Per-modality reconstruction terms are averaged so each modality contributes equally regardless of channel count.
  • Contrastive learning: The leave-one-out contrastive objective compares each modality embedding with the normalized mean of other present modalities and excludes unavailable modalities from its target.This formulation is designed to tolerate sample-level modality missingness.

Architecture ablations

The ablation isolates optimizer, positional-embedding, backbone, architecture, and objective changes while holding the pretraining corpus and token size fixed. SleepFM-2 combines one-second tokenization, channel-aware representations, a modern backbone, and both reconstruction and contrastive objectives.

  • Ablation design: The ablation chain adds one coherent change at a time while holding the corpus and token size fixed, then separately compares pretraining objectives.This design separates model changes from data and temporal-resolution effects.
  • Optimizer: Replacing SGD with AdamW improves age R2 by +0.06 and mean disease C-index by +0.02, with roughly neutral mean staging F1macro.The paper treats this as baseline modernization rather than an architectural change.
  • Channel-region position embedding: Adding channel-region position embeddings leaves disease and staging essentially unchanged but reduces age R2 by 0.05 in the channel-pooled v1 architecture.The embedding becomes useful only after the encoder attends across channels directly.
  • Modern backbone: The modern two-stage LLaMA-style backbone preserves channel-aware tokens and reduces encoder parameters from 4.83M to 2.56M.Its design uses joint attention over the channel-by-time token grid before temporal summarization.
  • Modern backbone: Despite fewer parameters, the modern backbone improves mean disease C-index by +0.01, age R2 by +0.08, and mean staging F1macro by +0.01.The reported improvements accompany the FFN and projection changes that reduce parameter count.
  • Pretraining objectives: Joint reconstruction and contrastive pretraining preserves reconstruction-only staging while adding 0.014 disease C-index and 0.121 age R2 over contrastive-only training.The two objectives contribute complementary dense per-token and summary-level information.
  • Token size: One-second tokens leave disease prediction at 0.798 and improve sleep staging to 0.773, while enabling events too short for five-second tokens.Age R2 changes by a small reduction to 0.758 in this experiment.

Downstream adaptation framework

SleepFM-2 uses a frozen encoder to generate reusable embeddings, then trains lightweight task heads for different cohorts and prediction tasks. The framework supports pooled subject-level representations, sequence modeling, and optional demographic augmentation.

  • Frozen adaptation: The pretrained encoder is frozen for every downstream evaluation, and only a task head is trained on precomputed embeddings.This lets the same embeddings serve multiple heads while decoupling encoder and task-training costs.
  • Embedding resolution: Each subject’s night is split into non-overlapping five-minute chunks, producing four modality-level summary embeddings per chunk.The resulting subject tensor has shape (M, S, dproj), with M = 4 modality streams.
  • Task heads: LinearProbe mean-pools modalities and chunks before a linear layer, whereas LSTMHead learns modality attention and models the chunk sequence with a transformer and bidirectional LSTM.The heads provide simpler and more expressive alternatives for downstream prediction.
  • Demographic augmentation: Age, sex, and BMI can be concatenated with the pooled PSG representation, while matched demographics-only heads provide baselines.The demographic vector is processed through a two-layer MLP before concatenation.

Disease prediction

Disease prediction evaluates frozen SleepFM-2 embeddings against interpretable PSG features and demographic baselines using held-out, cross-fitted analyses. The framework also tests whether embeddings contain shared and feature-specific information beyond demographics.

  • Evaluation design: Disease heads are trained on held-out cohort splits, while reported concordance and significance tests use complete held-out risk sets and six-year right-censored follow-up.The fitting loss uses minibatch risk sets only as an approximation.
  • Evaluation design: The primary significance test evaluates the standardized with-demographics risk score in a marginal Cox proportional-hazards model with Bonferroni correction.A parallel PSG-only head supplies the score used for significance testing.
  • Interpretable baseline: The interpretable baseline contains 480 validated PSG features spanning spectral EEG, microstructure, coupling, hypnodensity, cardiac, respiratory, and macro-sleep measurements.The six named physiological domains account for 469 features, with 11 additional macro summaries.
  • Interpretable baseline: Interpretable feature-bank arms use matched nonlinear multilabel Cox heads across all 939 phenotypes, including demographics-only and demographics-augmented variants.This keeps the modeling head consistent across feature-based comparisons.
  • Shared disease axis: The shared disease axis is the leading principal component of held-out phenotype risk scores after residualization and re-standardization on age, sex, and BMI.It describes the held-out score matrix rather than transporting a fitted model.
  • Feature recovery: Embedding features are tested by ridge-regressing each of the 480 interpretable measurements on the frozen 512-dimensional pooled embedding against a demographics-only probe.The comparison uses identical subjects and folds for both probes.
  • Condition-specific information: Condition-specific information is tested with nested Cox models comparing demographics, demographics plus the shared axis, and demographics plus the axis and condition-specific risk score.All models are cross-fitted over five folds with preprocessing learned only on training folds.

Sleep scoring

SleepFM-2 treats sleep scoring as dense sequence labeling across staging, arousals, limb movements, and respiratory events. One-second frozen embeddings support per-timestep heads, rare-event sampling, and comparisons with expert consensus.

  • Task scope: Sleep scoring covers five-class staging, binary cortical arousal detection, binary limb movement detection, and four-class respiratory event detection.All four tasks predict labels at every valid timestep from one-second frozen-encoder embeddings.
  • Task representation: Sleep-scoring heads use per-channel, one-second embeddings rather than five-minute subject-level summaries.Respiratory prediction additionally conditions on pretrained arousal and staging outputs through a cascade.
  • Training: Training uses masked, class-weighted cross-entropy, with event-centered sampling for sparse arousal, limb, and respiratory events.Respiratory sampling is further stratified by event class with fixed per-class quotas.
  • Expert comparison: Arousal and limb movement performance is evaluated on WSC recordings scored independently by nine technicians, while respiratory events use DREEM recordings scored by five raters.These datasets provide independent expert annotations for event-level comparison.
  • Expert comparison: Models and experts are compared against the same leave-one-out pseudo-consensus, formed from the remaining scorers for each held-out scorer.This aligns the model and expert evaluation target.

Representation analysis

Representation analysis tests what SleepFM-2 encodes directly in frozen embeddings, beyond task-head performance. The representation separates sleep stages and microstructure events, while additional analyses map feature coverage and concept entanglement.

  • Concept probes: Subject-level LinearProbe analyses use held-out AUROC with subject-disjoint splits, while continuous age is reported as rank/C-index encoding strength.Demographic probes provide context and are not comparable to subject-level age R2.
  • Sleep-geometry atlas: LDA projects deep-layer embeddings into two dimensions to visualize five-stage organization and overlay microstructure-event tokens in the same space.The atlas includes cortical arousals, N2 spindles, N2/N3 slow waves and REM eye movements.
  • Sparse-autoencoder feature taxonomy: Sparse autoencoders are evaluated across seven read-out depths and expansion ratios from E ∈ {1, 2, 4, 8, 16, 32, 64}, alongside raw embeddings.Features are attributed to macro, microstructure, demographics, pathology and disease domains.
  • Feature taxonomy and entanglement: Domain-level feature analysis distinguishes monosemantic features from entangled features spanning multiple concept groups, while reporting coverage, active-but-unaligned and silent features.Concept-direction cosine similarity measures entanglement strength from 0 for independent to 1 for collinear directions, with sign indicating co-activation or mutual exclusion.
  • Concept probes: 0.94 AUROC separates five sleep stages in frozen embeddings, while cortical arousals reach 0.92 and spindles, slow waves and eye movements reach 0.90–0.92.Separability increases with encoder depth.

Wakeful EEG and BCI transfer

The frozen SleepFM-2 encoder is evaluated on wakeful EEG and brain–computer-interface data using dataset-specific read-out depths and lightweight probes. The comparison contrasts pretrained frozen embeddings with end-to-end supervised training from random initialization.

  • Wakeful EEG: 23 wakeful EEG datasets are preprocessed into fixed-length 128-Hz windows and mapped to the pretrained channel-region vocabulary.Signals are normalized to match the scale used during pretraining.
  • Read-out strategy: Clinical, pathological, cognitive, affective and phenotyping datasets use Stage-2 pooled outputs, while motor-imagery datasets use flattened input-projection tokens.The read-out preserves the full time-patch sequence for the Stage-2 pathway and spatial layout for motor imagery.
  • Comparison: SleepFM-2 uses frozen embeddings with an L2-regularized LinearProbe, whereas E2E-Supervised trains the same architecture end to end from random initialization.The frozen condition selects regularization by inner cross-validation; the supervised comparator uses AdamW with separate encoder and head learning rates.

Wearable EEG and PPG sleep-staging transfer

SleepFM-2 transfers from multimodal PSG to headband, ear-EEG, PPG and wrist accelerometry through channel mapping and surrogate physiological signals. Frozen embeddings are used for sleep staging across six accelerometry cohorts and for UK Biobank disease prediction.

  • Wearable EEG and PPG: Wearable signals are harmonized to five-class Wake, N1, N2, N3 and REM labels, with invalid epochs excluded and channels mapped by anatomy.Missing device channels receive PAD indices and are masked; no channel is duplicated and no bipolar montage is reconstructed.
  • Wearable sleep staging: SleepFM-2 encodes 300-second windows as 1-second patches, mean-pools 30 patches per scoring epoch and applies a two-layer bidirectional LSTM head to fixed contexts.The comparator uses the same architecture but trains the encoder end to end from random initialization.
  • Wrist accelerometry: Respiratory and cardiac surrogate channels derived from three accelerometer axes are routed through the encoder’s existing PSG vocabulary without a student encoder or cross-modal alignment objective.The respiratory surrogate uses 0.1–0.6 Hz filtering, while the cardiac surrogate uses 3.5–14 Hz filtering followed by 0.5–3.5 Hz re-filtering.
  • UK Biobank disease prediction: For staging, 30 one-second embeddings form each 30-second representation; for UK Biobank, cardiac and respiratory embeddings are concatenated into one 256-dimensional vector per 300-second window.The UK Biobank head aggregates window embeddings with attention pooling, a transformer layer and bidirectional LSTMs before outcome-specific log-hazard prediction.
  • Wrist accelerometry: Accelerometry transfer evaluates five-class 30-second sleep staging across TBI, DREAMT, STAGES, Newcastle, Amazfit and SleepAccel, totaling approximately 453 subjects after quality control.The six cohorts include four internal cohorts and two external cohorts, with subject-level five-fold evaluation.
  • UK Biobank disease prediction: UK Biobank scoring uses the original evaluation code on 5,192 test participants, with participant bootstrapping and Benjamini–Hochberg correction across 101 outcomes with at least 50 events.The encoder remains frozen throughout training.

BioSerenity endpoints

The BioSerenity analysis evaluates frozen SleepFM-2 embeddings against demographic and tabular baselines on symptom and microstructure endpoints. Derived symptom scores use prespecified questionnaire-item rules and endpoint-specific labeled test sets.

  • Cohort and endpoints: 66,704 BioSerenity subjects are used for training, 1,295 for validation and 6,617 for testing, with 74,028 carrying all three microstructure feature sets.Endpoint-specific effective sample sizes vary according to label availability.
  • Symptom-score construction: Insomnia and Fatigue sum three questionnaire items, Parasomnia sums five, and Somatic combines eleven binary with five ordinal items before threshold binarization.Subjects missing any component item are excluded from the corresponding score.
  • Model comparison: SleepFM-2 uses an LSTMHead on frozen embeddings, while tabular arms use nonlinear DemoOnlyMLP heads with training-split median imputation and standardization.The tabular baselines are not restricted to linear combinations of their features.

Data availability

The study combines public, credentialed, proprietary and evaluation-only datasets, with access governed by repository rules, data-use agreements or institutional approvals.

  • The study uses 26 polysomnography cohorts and one wrist-accelerometry cohort for pretraining and evaluation, plus wearable and non-sleep EEG datasets for evaluation.
  • Twelve cohorts are publicly available through the National Sleep Research Resource after registration and acceptance of its data-use agreement.
  • The Human Sleep Project datasets require credentialed access, training and a data-use agreement.
  • UK Biobank accelerometry and linked outcomes are available to bona fide researchers through its application process and cannot be redistributed.
  • Most clinical PSG cohorts are available only on reasonable request, subject to institutional data-sharing agreements and local ethics approval.
  • BioSerenity recordings, which comprise most pretraining data and all subjective-sleep analyses, are proprietary and cannot be shared.

A Supplementary Material

The supplementary material documents cohort splits, model configurations, ablation results, phenotype criteria and shared latent structure, while noting that the structure analysis is coarser than the main-text analysis.

  • Data and splits: Supplementary Table 1 documents 235,865 pretraining recordings, frozen-encoder downstream evaluation, participant-level splits and held-out validation and test recordings.
  • Model comparisons: The SleepFM-1→SleepFM-2 ablation chain compares cumulative architectural changes and pretraining objectives, including contrastive, masked-autoencoding and joint losses.
  • Model comparisons: The supplementary performance table reports disease Harrell’s C-index, sleep-staging F1macro and age linear-probe R2 across the ablation chain.
  • Representation analysis: SleepFM-2’s shared latent structure keeps sleep stages, microstructure, demographics and disease broadly distinct, with mean within-group cosine similarity 0.20 versus 0.04 across groups.
  • Representation analysis: The shared-structure analysis is coarser than the main-text phecode analysis and measures organization in the encoder rather than in downstream models.
Loading 2609.06849v1…