Source-linked AI summary

Capturing Cardiac Cyclicity through Phase-Equivariant Self-Supervised Learning

Blaise Delaney, Dominic Dootson, Juan Jose Juan Castella, Salil Patel, Andrew Pfaff, Yuji Xing, Jonny Hancox, Karin Sevegnani

arXiv:2608.21147v1cs.LG

TL;DR

Self-supervised ECG learning must preserve cardiac timing and morphology without being dominated by nuisance signals, while standard latent prediction leaves phase organisation unspecified. WINDER adds a fixed phase-equivariant transport objective to LeJEPA, using an analytic harmonic action without transport parameters. On PTB-XL, it improves diagnostic performance over a matched control and reaches the broad range of larger self-supervised ECG models with 1.2 M deployed parameters.

  • Problem

    Self-supervised ECG objectives need to preserve cardiac-cycle information without overemphasising sensor artefacts, and standard latent prediction does not specify consistent phase organisation across beats.

  • Method

    WINDER extends LeJEPA with a transport objective and a fixed, closed-form harmonic action that encourages latent representations to follow cardiac phase without learned transport parameters.

  • Results

    Diagnostic performance improves over a matched transport-free control, while WINDER reaches the broad performance range of larger self-supervised ECG models with 1.2 M deployed parameters.

  • Takeaways & Limitations

    The study supports explicitly encoding cardiac-phase symmetry as a parameter-efficient way to organise diagnostically useful ECG representations around a measurable physiological quantity.

  • Takeaways & Limitations

    Evidence comes from a single corpus and simple encoder under a frozen linear readout, so broader corpora and stronger backbones remain future evaluations.

Abstract

from arXiv · show

The cyclic structure of physiological processes offers a natural prior for self-supervised representation learning, and the cardiac cycle provides a particularly well-defined setting in which to exploit it. We derive a phase-equivariant self-supervised objective and introduce Winder, a joint-embedding architecture that organises representations into phase-invariant coordinates and phase-rotating harmonic subspaces. Its transport operator is fixed and closed-form, derived from the cycle's geometry rather than learned, and adds no parameters. Evaluated on PTB-XL under a frozen linear-probe protocol, Winder attains diagnostic accuracy within the range reported by state-of-the-art self-supervised methods at a ~1 M parameter footprint, while exhibiting phase-equivariant latent geometry. These findings demonstrate that explicitly encoding cardiac-phase symmetry can preserve diagnostically useful information while yielding a latent geometry that is legible, parameter-efficient, and directly tied to a measurable physiological quantity.

1 Introduction

Cardiac-cycle structure motivates a phase-aware latent-prediction objective that avoids overemphasising sensor nuisance while consistently organising representations across beats. WINDER adds phase transport to LeJEPA and achieves competitive diagnostic performance with a compact model.

  • Physiological motivation: ECG records recurrent cardiac depolarisation and repolarisation whose timing and morphology can be aligned by mapping successive R–R intervals to a common phase coordinate.This phase coordinate preserves within-cycle ordering while normalising beat duration.
  • Self-supervised learning: Latent prediction can avoid explicitly rewarding reconstruction of motion, electrode artefacts, and other sensor-level nuisance signals.Unlike input-space reconstruction, joint-embedding objectives score predictions against latent signal states rather than raw samples.
  • Problem: Standard latent-prediction objectives leave cardiac-phase organisation unspecified, so representations may encode phase inconsistently across beats.The resulting gap motivates an explicit phase-equivariant constraint.
  • Contribution: WINDER extends LeJEPA with a transport objective that penalises deviations from phase equivariance and encourages cardiac-phase-ordered latent trajectories.The added objective complements LeJEPA’s anti-collapse regularisation.
  • Results: Diagnostic performance improves over a matched transport-free control, while WINDER reaches the broad performance range of larger self-supervised ECG models with only 1.2 M deployed parameters.The model also produces phase-ordered latent geometry.

2 Related Work

Prior work addresses collapse, temporal representation learning, cyclic signals, and equivariance, but does not establish a derived cardiac-phase action on learned ECG latents. WINDER targets this unoccupied combination by imposing phase structure directly in latent space.

  • Predictive self-supervision: LeJEPA combines predictive learning with SIGReg, which tests Gaussian isotropy along randomly sampled projection directions to prevent representational collapse.Its temporal extension applies the same regularisation to next-embedding prediction.
  • Predictive self-supervision: SIGReg-based methods and related JEPA variants revise target geometry, test statistics, or temporal pooling, while their regularised objectives generally remain invariant to time-index permutations.The conditions supporting their guarantees are still being characterised.
  • Cyclic biosignal learning: Phase-aware ECG and periodic-biosignal methods predict or accommodate cyclicity, but do not impose cardiac phase as a symmetry of the latent space.CardioState-JEPA treats phase as an auxiliary supervised target, whereas Sonata models oscillatory modes in other biosignals.
  • Equivariant representation learning: Equivariant representation learning constrains latent variables to transform under a known group action, often separating invariant factors from transforming coordinates.Related examples span images, generic time series, audio, and health applications.
  • Equivariant representation learning: ECG methods such as LVCG and Nef-Net model electrode-view geometry, but view invariance differs from controlled equivariance to cardiac phase.Other health approaches likewise use cyclic shifts or phase invariance rather than physiological phase advancement.
  • Research gap: A derived cardiac-phase action on a learned ECG latent remains unoccupied in prior work, which WINDER is designed to address.Its rotation acts on a physiologically derived cardiac-phase coordinate rather than the raw time axis.

3 Theory

WINDER reparameterizes cardiac recordings by phase and requires latent representations to transform consistently under phase translations. Its fixed harmonic decomposition combines invariant coordinates with rotating two-dimensional channels, while prediction, transport, and spectral objectives regulate the representation.

  • Phase representation: Cardiac recordings are reparameterized from time to phase on the circular domain S1, inducing a phase-translation group isomorphic to U(1).The phase map is based on detected R-peak timestamps, and the group acts on phase-parameterized 12-lead ECG signals.
  • Phase representation: WINDER requires phase translation before encoding to equal the corresponding latent transformation after encoding: fθ(T∆x) = R∆fθ(x).This is the defining phase-equivariance condition for the learned representation.
  • Latent geometry: The canonical real representations comprise a trivial zero-frequency component and two-dimensional rotation blocks for positive harmonic orders of U(1).An orthogonal basis change Q expresses learned latents in the canonical irreducible-representation basis; a learned Q is possible, but the selected architecture fixes the canonical basis and therefore uses no learnable transport parameters.
  • Latent geometry: The latent action decomposes into invariant and harmonic subspaces, with K0 coordinates fixed and cn two-dimensional channels rotating at integer phase rates n∆.The multiplicity cn specifies the number of independent feature channels assigned to harmonic n.
  • Learning objective: WINDER combines latent prediction, phase-equivariant transport, and spectral regularisation to preserve predictive content, constrain phase-dependent directions, and prevent dimensional collapse.The transport term constrains unit-normalised latent directions, while prediction and spectral objectives control latent magnitudes and distribution.
  • Learning objective: The transport objective alone can be satisfied by concentrating representations in the invariant block, so it does not distinguish useful phase equivariance from trivial phase invariance.This is a phase-blind failure mode because the invariant block is unchanged by R∆.

4 Method

WINDER is implemented as a causal joint-embedding architecture for 12-lead ECG, using phase-indexed transport to constrain projected representations according to cardiac phase. The method combines a compact encoder–projector deployment path with a training-only causal predictor and fixed harmonic transport structure.

  • Dataset and evaluation: PTB-XL provides 10-second, 12-lead ECG recordings sampled at 500 Hz for phase-aware representation learning and diagnostic evaluation.The corpus contains 21,799 recordings from 18,869 patients, with cardiologist-assigned diagnostic, form, and rhythm statements.
  • Dataset and evaluation: The study evaluates frozen latent representations on PTB-XL diagnostic subclasses using 23-subclass and five-group macro-average metrics.The five-group metric follows PTB-XL’s hierarchy of normal ECG, myocardial infarction, ST/T change, conduction disturbance, and hypertrophy.
  • Dataset and evaluation: The confirmatory protocol pre-trains on patient-disjoint folds 1–9 and evaluates a frozen linear probe on sealed fold 10.The confirmatory pre-training set contains 19,601 recordings from 16,965 patients, while fold 10 contains 2,198 recordings from 1,904 patients.
  • Input representation: Each 10-second recording becomes 125 non-overlapping 80 ms tokens, while cardiac phase is estimated by interpolating between native-resolution R-peaks.The encoder operates on signals resampled to 100 Hz, and tokens are assigned the phase at each patch centre; samples outside detected peak boundaries are excluded from phase-dependent analyses.
  • Encoder and projector: The causal encoder uses patch MLP processing and two residual convolutional context blocks, followed by a shared token-wise projector producing 256-dimensional projected representations.The encoder processes 12 × 8 patches through dimensions 96 → 512 → 256, while the projector maps 256 → 512 → 256 without mixing information across time.
  • Phase transport: The projected space allocates K0 = 4 invariant dimensions and 252 harmonic dimensions, with each harmonic coordinate pair rotated by n∆.Harmonics n = 1, . . . , 10 use multiplicities (24, 24, 20, 16, 12, 10, 8, 6, 4, 2), and the transport operator has zero learnable parameters.
  • Training and deployment: The encoder and projector form a 1,232,384-parameter deployed path, while a 3,227,084-parameter causal Transformer predictor is used only during self-supervised training.The predictor raises the optimised total to 4,459,468 parameters and is discarded after pre-training; downstream probes use mean-pooled projected tokens with defined phase.
  • Training and deployment: Training minimises a three-term projector-output objective, with λSIG = 0.15 and λtrans = 1, while the transport-free control sets λtrans = 0.Prediction targets occur after independently sampled causal cutoffs, and inputs after each cutoff are replaced by a learned mask token.

5 Results

A matched ablation shows that phase transport improves diagnostic accuracy, latent organisation, and robustness at identical inference cost, while mechanism tests support cardiac-clock use and phase-structured geometry.

  • Ablation design: The matched WINDER-control ablation differs only in transport, with λtrans = 1.0 for WINDER and λtrans = 0.0 for the control.Both arms share data, architecture, seeds, schedule, and evaluation protocol.
  • Diagnostic accuracy: +0.0774 to +0.0910 superclass gap across the full 12-cell grid, with no overlapping matched cross-arm confidence intervals.The two seeds agree, and the largest WINDER between-seed spread is 0.0054.
  • Mechanism validity: WINDER attains mean gain ḡ = 0.8182 and gain fraction gf = 0.8786, passing G1, whereas the control has gf = −0.0744 and fails.The test compares the true cardiac clock with shuffled phase labels; the shuffled-magnitude condition discriminates the arms.
  • Latent geometry: 1.8941 versus 0.3817 amplitude indicates a roughly five-fold stronger phase-locked fundamental component for WINDER than the control.Mean coherence is 0.9927 for WINDER versus 0.3159 for the control; coherence measures common-plane rotation, not loop order.
  • Latent geometry: WINDER forms a phase-ordered annulus and closed circular trajectories, while the control lacks the corresponding ordering and closure.The UMAP projection receives no phase, arm, or diagnosis information and colours points by phase only afterward; the geometry is qualitative and uses folds 1–9.
  • Anomaly detection: 80 ms is the finest anomaly-detection delay expressible by the 80 ms token grid, not a measured bound on response speed.The reported WINDER median is 80 ms, and shorter patches would be required to resolve faster detection.

6 Discussion

WINDER uses a fixed harmonic phase action and transport objective to organise cardiac representations, with diagnostics assessing whether the learned geometry follows that action. Controlled comparisons show performance and robustness benefits, while the evidence remains bounded by the study’s corpus, backbone, evaluation protocol, and synthetic anomaly tests.

  • Phase geometry: Fixed harmonic rotations provide a soft physiological prior, encouraging phase-predictable latent variation while guaranteeing cycle closure through R2π = I.Transport gain and latent-geometry measurements assess adherence to the prescribed action.
  • Empirical effects: Phase transport improves diagnostic accessibility and robustness to lead loss relative to a matched control that disables only the transport loss.The comparison attributes the isolated difference between models to phase transport.
  • Monitoring: The declared geometry makes phase organisation and transport violations directly quantifiable as a monitorable contract for a compact causal encoder.The proposed view is motivated for continuous monitoring at acquisition, including ambulatory telemetry and bedside rhythm surveillance.
  • Scope: Phase-aware anomaly scores were evaluated only on synthetic perturbations and do not establish performance in a deployed clinical system.The passage frames continuous monitoring as a supported possibility rather than a validated clinical result.
  • Scope: Evidence comes from a single corpus, a deliberately simple encoder, and a frozen linear-readout comparison rather than a common-protocol leaderboard.The authors call for broader corpora, stronger backbones, feature analysis, and multi-step forecasts.

7 Conclusions

WINDER introduces a phase-equivariant objective that represents cardiac phase through a fixed, closed-form latent action with invariant and harmonic subspaces. The transport objective encourages adherence to this action, which is evaluated using transport gain and latent-geometry diagnostics.

  • WINDER represents cardiac phase with a fixed, closed-form latent action whose invariant and phase-rotating harmonic subspaces require no learnable transport parameters.
  • The transport objective encourages learned representations to follow the prescribed phase action.
  • Transport gain and latent-geometry diagnostics measure whether the learned representation follows the prescribed action.

A Data protocol

The data protocol specifies preprocessing, normalization, phase assignment, and quality-control handling for the PTB-XL corpus. Training uses patch-centre phase, while causal phase is reserved for reporting descriptors that cannot read beyond each token’s support.

  • Protocol: The appendix documents the preprocessing chain, phase-clock construction, normalization convention, and pretraining-pool composition.
  • Normalization: Fixed per-lead normalization fits one mean and standard deviation on 19,601 decimated training-fold recordings and applies them unchanged to every recording.Unlike per-beat or per-recording normalization, this preserves between-recording amplitude variation within each lead.
  • Corpus: All 19,601 fold-1–9 recordings are used for pretraining, while 222 recordings failing phase quality control are retained because phase is not consumed by the model.The retained failures represent 1.0% of the corpus.
  • Phase assignment: Each 80 ms token summarizes an eight-sample decimated-signal patch, requiring a single scalar phase for the whole patch.
  • Phase assignment: Training assigns each token the phase at its patch centre, whereas reporting-only causal descriptors use the phase at the token’s last sample.The causal convention prevents descriptors from reading beyond a token’s support and never enters training.
  • Phase assignment: A training-time causal boundary would impose a fixed 35 ms offset whose phase error increases with harmonic order, reaching 0.39 rad at n = 7.The highest retained harmonic determines timestamp tolerance because harmonic phase advances n times faster.

C Training specification

The training specification fixes projection sampling and quadrature choices for the reported runs, and compares alternative phase frames that were not adopted. The selected configuration retains raw token representations and disables record-level SIGReg.

  • LSIG: All reported runs evaluate LSIG on raw projector tokens using M = 256 freshly sampled random projection directions per training step.
  • Quadrature: Equation (2)’s characteristic-function integral uses J = 17 trapezoidal knots over [0, 3], exploiting integrand evenness to represent [−3, 3].The shorter interval provides finer resolution near the origin while Gaussian weighting reduces the importance of |t| > 3.
  • Frame comparison: Canonical phase-zero demodulation produced transport gains from −0.38 to −0.17 and reduced effective rank to 2.9–11.2.
  • Frame comparison: Record-canonical templates yielded zero transport gain because representations concentrated in the nonrotating invariant block.
  • Selected configuration: Reported runs therefore use the raw token frame with record-level SIGReg disabled.

C.2 Augmentation provenance

The augmentation stack was motivated by related time-series and ECG work, then carried forward after a preliminary single-seed comparison, while its individual effects remain unisolated.

  • Augmentation provenance: Related work motivated signal augmentation, but did not validate the specific six-transform stack used here.The cited precedent includes Reverso and ER-JEPA, which used augmentation in related settings.
  • Augmentation provenance: The stack was selected through preliminary architecture scouting using one seed, 5,000 training steps, 6,000 records, and fold 9 evaluation.
  • Augmentation provenance: The augmentation settings are documented in Table 6, while preprocessing and cohort composition are documented in Tables 3–5.
  • Augmentation provenance: 0.8589 macro-AUROC was achieved with the full augmentation stack versus 0.8513 for the control arm.
  • Augmentation provenance: The preliminary comparison supports carrying the stack forward but does not estimate uncertainty or isolate individual augmentation contributions.

C.4 Four training runs

Four pre-training runs used matched transport-enabled and transport-free arms with fixed optimisation settings, independent randomness streams, checkpointing, and matched GPU execution.

  • Four training runs: The experiment used two arms and two seeds, with the transport-enabled arm setting λtrans = 1 and the control assigning zero transport-loss weight.
  • Four training runs: Table 7 lists the fixed optimiser and schedule settings used for all reported pre-training runs.
  • Four training runs: All remaining model, data, and optimisation settings were identical across the four runs.
  • Four training runs: Randomness was separated into independent streams for initialisation, masking, SIGReg projections, augmentation, and data order, with stream states stored in checkpoints.
  • Four training runs: Checkpoints were saved every 2,500 steps and at the end of training.
  • Four training runs: Training ran on one NVIDIA A100 40 GB GPU in two matched concurrent pairs, using approximately 11.7 GB peak memory per pair.

D Positioning within PTB-XL self-supervised learning

WINDER is positioned as a compact phase-equivariant ECG representation learner rather than a directly comparable leaderboard entry, because published PTB-XL evaluations use differing protocols and readouts.

  • Positioning within PTB-XL self-supervised learning: The curated comparison covers PTB-XL methods reporting diagnostic performance with sufficient information on pre-training data, model scale, and downstream evaluation protocol.
  • Positioning within PTB-XL self-supervised learning: Published PTB-XL results do not share a common evaluation protocol, so the comparison provides methodological and empirical reference points rather than a leaderboard.
  • Positioning within PTB-XL self-supervised learning: LeNEPA is the nearest methodological comparison: both combine causal latent prediction with SIGReg, but only WINDER imposes cardiac-phase meaning through equivariant transport.
  • Positioning within PTB-XL self-supervised learning: Other reference methods differ in scale, pre-training corpora, task definitions, or fold specification, including S4-JEPA, Weimann JEPA, ECG-CPC, and Poly-Window.
  • Positioning within PTB-XL self-supervised learning: The comparison does not support a head-to-head accuracy claim because Weimann uses learned cross-attention pooling whereas WINDER uses parameter-free mean pooling.
  • Positioning within PTB-XL self-supervised learning: WINDER differs from CardioState-JEPA by using ECG alone and imposing phase as a fixed latent-space group action rather than predicting it.
  • Positioning within PTB-XL self-supervised learning: WINDER uses a 1.2M-parameter model, 19,601 PTB-XL pre-training records, parameter-free mean pooling, and a linear classifier.
  • Positioning within PTB-XL self-supervised learning: WINDER’s diagnostic performance lies in the same broad range as selected self-supervised models despite differing evaluation settings.
Loading 2608.21147v1…