Source-linked AI summary
Behavioral Latency as Weak Event-Time Supervision for EEG Reaction-Time Decoding
Anuar Aimoldin, Ayana Mussabayeva, Yedige Mussabayev, Xue Liu, Kun Zhang
TL;DR
EEG reaction-time decoding typically treats latency as a scalar label, potentially discarding when response-relevant dynamics occur. This paper models RT as weak supervision for a posterior over event times and finds consistent held-out gains from distributional supervision, while diagnostics reveal partial temporal localization and a remaining equivariance gap.
Problem
Fixed-window scalar RT prediction can discard temporal structure and may succeed without representing when response-relevant EEG dynamics occur.
Method
The model treats behavioral RT as weak event-time supervision, learns p(t_event | X), and uses its posterior expectation for scalar prediction under controlled EEG evaluations.
Results
Distributional event-time supervision consistently improves held-out RT prediction over scalar and temporal-readout controls across seeds and backbones, with posterior diagnostics revealing partial crop-relative localization.
Takeaways & Limitations
Latency labels can preserve temporal meaning while remaining compatible with standard behavioral metrics and enabling direct tests of temporal evidence.
Takeaways & Limitations
Generality beyond the HBN-EEG contrast change detection task remains untested, and sensitivity remains below full crop-relative localization.
Abstract
from arXiv · showhide
Single-trial EEG analyses are often organized around events and latencies, yet EEG-based reaction-time (RT) prediction is posed as scalar regression on a fixed stimulus-locked window. RT is treated as a window-level label rather than timing evidence about response-relevant dynamics. Here we reformulate trial-wise RT decoding as event-time posterior modeling. Instead of predicting RT directly, the model estimates a posterior over response-relevant event times, $p(t_{\mathrm{event}}\mid X)$, and uses its mean as the RT estimate. This treats behavioral latency as a weak observation of latent response-relevant timing. We evaluate this formulation on the Healthy Brain Network contrast change detection EEG task under a subject-disjoint, release-separated protocol. Across five seeds, distributional event-time supervision consistently improves held-out RT prediction relative to scalar regression and temporal-readout controls. Controlled objective comparisons isolate supervision of the event-time distribution, rather than expectation-based readout alone, as the source of this gain. Architecture controls show that the effect persists across four temporal backbones and is not explained by model scale. Beyond point prediction, posterior geometry characterizes concentration, target alignment, and interval behavior, while observation-noise calibration separates latent concentration from predictive uncertainty over RT. Shifted-crop inference probes shortcut use versus temporal localization. Matched shift-jitter improves robustness, increases mean sensitivity, and moves predictions more often in the expected crop-relative direction. Sensitivity remains below ideal crop-relative localization, leaving a clear equivariance gap. Together, these results establish event-time posterior modeling as a probabilistic and interpretable formulation for linking single-trial EEG dynamics to behavioral timing.
1 Introduction
The paper reframes EEG reaction-time decoding as posterior modeling over response-relevant event times, using behavioral latency as weak timing supervision. Controlled comparisons test whether distributional supervision improves prediction and exposes temporal behavior beyond scalar accuracy.
- Scalar RT prediction can exploit individual response tendencies or stimulus-locked priors without representing when response-relevant dynamics occur.
- The model estimates p(t_event | X) and uses its posterior expectation as the scalar RT prediction.
- Event-time modeling treats button-press latency as noisy evidence about response-relevant dynamics shaped by sensory, decision, and motor processes.
- Controlled comparisons separate backbone capacity, temporal readout parameterization, scalar timing losses, and distributional event-time supervision.
- Distributional event-time supervision improves held-out prediction over scalar, temporal-readout, and soft-argmax controls across five seeds and four dense temporal backbones.
- Posterior geometry, observation-noise calibration, and shifted-crop diagnostics assess concentration, uncertainty, interval behavior, shortcut use, and temporal localization.
2 Related Work
Related work frames EEG latency targets as temporally structured outputs rather than ordinary window-level labels. The paper connects event-time supervision with distributional prediction and distinguishes its output-representation focus from input-representation approaches.
- EEG research studies response preparation, errors, movement onset, ERP timing, and transient events through temporally localized neural activity.
- Fixed-window scalar decoding is convenient for benchmark evaluation but can collapse temporal structure when the target is a latency.
- Reaction-time EEG work commonly uses scalar regression and normalized RMSE, while this paper retains that evaluation protocol but adds timing semantics.
- Ordered RT labels motivate smoothed event-time distributions, while time-to-event modeling supplies a probabilistic framework for event-time prediction and calibration.
- The paper’s contribution is an output-representation axis complementary to cross-subject, cross-session, and cross-dataset input-representation methods.
3 Dataset and Evaluation Protocol
The evaluation uses the HBN-EEG contrast change detection task with fixed stimulus-locked windows and release-separated, subject-disjoint splits. The protocol controls data support, comparisons, and diagnostics while restricting conclusions to modeled RTs.
- The CCD task provides 2 s stimulus-locked EEG windows with 128 channels and 200 samples, alongside RT from stimulus onset to button press.
- The benchmark fixes data splits, target support, input representation, and scalar evaluation while varying architectures and objectives.
- Training uses R1–R8, development uses R9–R10, and final holdout evaluation uses R11 with zero subject overlap across partitions.
- The analyzed sample excludes RTs outside 0.5–2.5 s, removing 2.07% of training, 2.90% of development, and 3.73% of holdout trials.
- Scalar performance is evaluated with normalized RMSE, while shifted-crop and posterior-only analyses are separate diagnostics that do not change main holdout reporting.
- The comparison includes scalar controls, output-supervision blocks, posterior-geometry diagnostics, shifted-crop inference, and matched shift-jitter intervention.
4 Scalar Regression Controls
The scalar-control analysis tests how far window-level RT supervision can perform using compact temporal regressors and external EEG architectures. ETR-CNN large is selected as the strongest scalar reference for event-time comparisons.
- Scalar baselines include compact task-specific temporal regressors and external EEG backbones under the same protocol.
- MSP-CNN summarizes coarse temporal segments with mean and max statistics to provide early, middle, and late activation cues for scalar RT.
- ETR-CNN produces learned temporal scores over the input window and reads them out for scalar RT prediction.
- ETR-CNN large is the strongest scalar baseline, with MSP-CNN and base ETR-CNN close behind.
- External EEG backbones remain below task-specific temporal controls, indicating that generic backbone capacity alone does not explain stronger scalar performance.
5 Event-Time Posterior Formulation
Event-time posterior modeling replaces direct scalar RT regression with distributional supervision over response-relevant times, while retaining posterior-mean RT readout. Controlled comparisons show gains beyond expectation-based readout, with robustness across objectives and temporal backbones.
- 5 Event-Time Posterior Formulation: Event-time models produce per-time logits and a posterior over response-relevant event times, then read out scalar RT as the posterior expectation.The formulation includes soft-target distribution matching and likelihood-based EventNLL objectives.
- 5.1 Primary Event-Time Segmentation Architecture: The ETS-U-Net maps EEG windows to dense temporal outputs using an encoder–decoder with multi-scale context and skip connections that preserve temporal resolution.Its segmentation design matches the event-time output space and supports controlled supervision comparisons.
- 5.2 Soft-Target Distribution Matching: Soft-target matching converts RT to crop-relative time and constructs a Gaussian target distribution whose bandwidth controls supervision smoothing rather than RT uncertainty.Cross-entropy is the primary matching objective, while Wasserstein distance provides a geometry-aware control.
- 5.2 Soft-Target Distribution Matching: The RT-only control supervises only posterior-mean RT, whereas CE and W1 supervise the full event-time distribution.This comparison isolates distributional supervision from expectation-based temporal readout.
- 5.3 Likelihood-Based Event-Time Objectives: Likelihood-based objectives treat observed RT as a noisy measurement of a latent response-relevant event time without constructing an explicit soft target.This provides a probabilistic alternative to Gaussian target construction within the same posterior framework.
- 5.4 Formulation Comparison and Robustness: Holdout τ-nRMSE was 0.8745–0.8778 for CE and likelihood-based variants, versus 0.8928 ± 0.0042 for ETR-CNN large.These results were obtained under a common readout-temperature procedure and show improvement beyond a strong scalar temporal-readout baseline.
- 5.4 Formulation Comparison and Robustness: Mixture EventNLL reduced τ-nRMSE from 0.8928 ± 0.0042 to 0.8745 ± 0.0053, corresponding to a 6.24 ms RMSE reduction and an 8.94 ms MAE reduction.The improvement was positive in all five matched-seed comparisons, and both subject-bootstrap confidence intervals excluded zero.
- 5.5 Architecture Robustness of the Supervision Effect: Across U-Net, dilated-convolution, multi-scale convolution, and attention-convolution backbones, CE and mixture EventNLL outperformed the RT-only soft-argmax control.The supervision advantage therefore replicated across four temporal architectures rather than depending on one backbone.
6 Posterior Readout and Diagnostics
The posterior mean provides a scalar RT readout, while posterior geometry, calibration, and shifted-crop tests reveal temporal concentration, uncertainty, and crop-relative behavior beyond scalar error.
- 6.1 Scalar Readout and Temperature Tuning: The posterior expectation converts event-time probabilities into an RT estimate, with the relative prediction shifted by the fixed window start for nRMSE comparison.Readout temperature is tuned on development data for scalar comparison, while probabilistic calibration and posterior concentration are evaluated separately.
- 6.2 Posterior Geometry Diagnostics: Event-time losses with similar scalar error produce different posterior geometry, including differences in width, target-aligned mass, mode–mean disagreement, and interval coverage.These metrics distinguish scalar accuracy, distributional scoring, temporal concentration, target alignment, and interval behavior.
- 6.2 Posterior Geometry Diagnostics: CE and mixture EventNLL provide the strongest scalar readouts, whereas EventNLL-family objectives are sharper and more target-concentrated but have low latent-posterior coverage.Wasserstein is closer to nominal coverage but weaker in scalar accuracy and near-target mass; the soft-argmax control has broad posteriors and the lowest Coverage MAE.
- 6.2 Posterior Geometry Diagnostics: Observation-noise calibration separates latent event-time concentration from predictive RT uncertainty and reduces holdout Coverage MAE to 0.005–0.008, with Coverage80 of 0.790–0.795.The EEG model, event-time posterior, and posterior-mean prediction remain fixed during calibration.
- 6.3 Shifted-Crop Shortcut-vs-Localization Diagnostic: Shifted-crop inference tests whether predictions move with the temporal frame, and matched shift-jitter improves shifted-crop accuracy, direction agreement, and sensitivity while sensitivity remains below 1.Wasserstein shows the strongest crop-relative sensitivity and direction agreement but weaker scalar and shifted-crop accuracy.
7 Discussion
The discussion frames the target representation as a substantive modeling choice: event-time objectives improve prediction and expose posterior behavior, but temporal equivariance and broader generality remain unresolved.
- 7 Discussion: Making behavioral latency explicit in the output space and loss supports stronger scalar readouts than the soft-argmax RT-loss control using the same posterior-mean readout.CE and EventNLL-family objectives form the strongest scalar-readout group in this study.
- 7 Discussion: Objective choice shapes posterior semantics: CE gives broad support, EventNLL models latent event time through an observation model, and Wasserstein is more localizer-like but weaker as a point predictor.These objectives therefore differ in both scalar accuracy and temporal behavior.
- 7 Discussion: Posterior width, target-aligned mass, mode–mean disagreement, likelihood, and coverage expose temporal behavior that scalar nRMSE discards.Similar point errors can correspond to broad conservative, sharp target-concentrated, or better-covered distributions.
- 7 Discussion: Shift-jitter training increases crop-relative sensitivity and direction across objectives, but sensitivity below 1 shows that localization remains partial.The residual gap is a quantitative target for models designed for temporal equivariance.
- 7 Discussion: Compact task-specific temporal models outperform the evaluated standardized and foundation-style backbones, while distributional objectives outperform RT-only posterior-mean training under matched controls.Within this benchmark, the controls identify formulation and objective design as sources of improvement beyond architecture capacity alone.
- 7 Discussion: The empirical scope is limited to the HBN-EEG contrast change detection task under a release-separated protocol, so broader generality remains to be tested.The shifted-crop analysis also leaves full crop-relative localization as an open target.
8 Conclusion
The paper reframes reaction time as weak supervision for event-time posteriors, improving RT decoding while exposing temporal structure and uncertainty beyond scalar error.
- Distributional event-time supervision consistently improved held-out RT prediction over direct scalar regression and an RT-only posterior-mean control.Five-seed comparisons and four temporal backbones attribute the advantage to supervised output formulation rather than model scale, expectation readout, or architecture.
- Posterior geometry distinguishes concentration, target alignment, interval behavior, and predictive uncertainty over observed RT.Observation-noise calibration separates latent event-time concentration from behavioral-RT predictive coverage.
- Shifted-crop analyses reveal partial crop-relative localization that scalar nRMSE alone cannot diagnose.Shift-jitter improves robustness and expected-direction movement, but sensitivity remains below ideal crop-relative localization.
- Latency labels can serve as structured supervision that preserves temporal meaning while remaining compatible with standard behavioral metrics.The paper proposes extending this formulation to other latency-defined EEG phenomena.
A Additional Readout, Calibration, and Shifted-Crop Details
Additional analyses specify how posterior readout, observation-noise calibration, and shifted-crop diagnostics are selected and interpreted across seeds.
- Calibration: Observation-noise calibration evaluates latent posterior coverage separately from predictive RT coverage after fixing the trained model and scalar readout.Calibrated kernel scale is selected on R9–R10 using Coverage MAE across central intervals from 50% to 90%.
- Readout temperature: Readout temperature is selected on the R9–R10 development split and applied unchanged to holdout predictions and posterior diagnostics.Different event-time objectives have different development-split nRMSE minima, indicating distinct posterior confidence scales.
- Shifted-crop diagnostics: Shift-jitter lowers mean shifted-crop relative nRMSE and increases sensitivity and direction across evaluated objectives.These seed-level changes move predictions toward crop-relative temporal localization, while absolute sensitivity remains below ideal.
B Architecture-Control Details
Architecture controls test whether shifted-crop behavior and posterior geometry persist across four dense temporal backbones rather than reflecting one model family or scale.
- Diagnostic comparisons retain common shifted-crop and posterior-geometry measures across dense temporal backbones.Posterior columns include a common distributional score, target-aligned mass, and interval coverage error.
- CE and mixture EventNLL consistently achieve lower shifted-crop error and more crop-responsive predictions than RT-only posterior-mean training across all four architectures.Higher sensitivity and direction scores indicate the replicated crop-response pattern.
- Absolute sensitivity remains below 1, indicating partial rather than fully crop-relative localization.The metric therefore quantifies an equivariance gap even when objective-level improvements replicate.
C Reproducibility Details
The reproducibility materials document configurations, shared experimental settings, architecture instantiation, stored outputs, and planned release artifacts for the controlled comparisons.
- YAML configurations specify the reported experiments, while experiment directories store training summaries, predictions, selected temperatures, and model summaries.The configuration repository is identified as the source for benchmark settings.
- The reproducibility summary covers data splits, input representation, optimization, augmentation, model selection, temperature tuning, and architecture hyperparameters.Objective-specific loss choices are described separately in the main text.
- External EEG architectures are instantiated from named Braindecode constructors and trained from scratch under a shared evaluation protocol.The protocol holds splits, target support, optimizer, augmentation, checkpoint selection, and holdout evaluation consistent.
- Planned release materials include preparation code, runners, model and loss code, seed summaries, diagnostics, aggregate tables, and figure-generation scripts.When raw HBN-EEG redistribution is restricted, the release will provide preparation code and expected split structure.