Source-linked AI summary

Neuro-Geospatial Modelling of EEG Affective States Using Literature-Informed Environmental Context

Utsav Poudel, Jagannath Aryal, Subramaniyaswamy Vairavasundaram

arXiv:2608.20807v1cs.AIcs.HCcs.LG

TL;DR

The paper addresses the lack of joint georeferencing between EEG affective-state benchmarks and environmental data. It uses a dual-tower EEG–environmental architecture with literature-informed priors and controlled evaluations, achieving 76.2% accuracy in Astana and 72.8% after Singapore environmental substitution, while remaining limited to technical feasibility rather than observed or causal exposure–affect inference.

  • Problem

    EEG affective-state benchmarks rarely include participant-level spatial or temporal exposure measurements, limiting direct integration with environmental context.

  • Method

    A dual-tower model combines an EEG-Conformer with a graph neural network environmental encoder, using literature-informed priors and label-independent evaluation after training.

  • Results

    76.2% accuracy in Astana fell to 72.8% when only the environmental tower was refitted to Singapore’s independently modelled distribution.

  • Takeaways & Limitations

    The framework demonstrates technical feasibility for combining unco-registered EEG and spatial environmental context and offers a reusable computational template for future jointly collected studies.

  • Takeaways & Limitations

    Because environmental context was constructed from literature priors rather than measured alongside EEG, the results do not establish observed or causal exposure–affect coupling.

Abstract

from arXiv · show

Environmental exposures such as air pollution and greenness have been associated with affective and cognitive outcomes, but EEG and environmental datasets are rarely jointly georeferenced. We investigate whether literature-informed environmental priors can serve as an auxiliary geospatial modality for EEG-based affective-state classification when individual-level exposure data are unavailable. We combine 30-channel EEG from the EAV benchmark (42 participants, aged 20-30 years) with environmental representations derived from OpenAQ, Sentinel-2, Sentinel-5P, and OpenStreetMap data for Astana. A dual-tower architecture combines EEG-Conformer representations with a graph-based environmental encoder. Because the datasets are not co-registered, environmental context is treated as a literature-informed prior rather than measured exposure. Subject-level repeated splits, permutation and label-shuffling controls, dose-response reversal, and domain-shift experiments distinguish architecture-level gains from prior-dependent gains. The multimodal model achieves 76.2% accuracy versus 67.4% for EEG alone. Controls disrupting environmental-label structure retain part of this gain, indicating that the improvement is not attributable solely to environmental information. Replacing the Astana environmental distribution with an independently modeled Singapore distribution reduces accuracy to 72.8%. These findings demonstrate technical feasibility but do not establish an observed or causal exposure-affect association. The study provides a framework for future jointly collected mobile EEG-environment studies. Implementation: https://github.com/r11up/geo-cog

1 Background

The study frames EEG affective-state research as disconnected from real-world environmental context and proposes a literature-informed, spatially structured prior to computationally link the domains without claiming measured exposure–affect relationships.

  • Research gap: EEG affective-state research is largely laboratory-bound, while few studies directly combine EEG with spatially distributed environmental context.EEG offers millisecond-scale sensitivity to affective fluctuations, but benchmark recordings generally lack simultaneous outdoor location or exposure measurements.
  • Research gap: EAV contains laboratory EEG from Astana without participant-level trajectories, outdoor sensor readings, or timestamped exposure measurements.Environmental vectors therefore represent city-wide distributions rather than each participant’s exposure during recording.
  • Approach: Environmental context is generated as a literature-informed computational prior rather than measured exposure, so the framework evaluates model behaviour rather than causal environmental effects.The environmental signal is sampled from city-wide distributions and conditioned during training using literature-derived dose-response relationships.
  • Approach: The proposed framework combines an EEG-Conformer with a graph neural network over a 100-cell environmental grid using bidirectional cross-modal attention.The environmental tower encodes spatial features including pollution, vegetation, and urban morphology, while the two towers are trained jointly.
  • Evaluation: The reusable two-city design tests whether the environmental tower can be adapted to an independently modelled city while keeping the EEG data framework fixed.Astana supplies the primary EEG setting, while Singapore provides an environmental-domain shift evaluation only.
  • Evaluation: The study uses subject-level repeated splits and controls designed to separate architecture-level gains from benefits dependent on the imposed environmental-label prior.Evaluation includes label-independent validation and testing, environmental-domain substitution, and component-isolating configurations.

3 Results

The multimodal model outperformed EEG-only classification under label-independent inference, while controls showed that the gain combined architecture-level and literature-informed prior effects. Performance declined under environmental domain shift and coarser spatial resolution, and feature attributions described reliance on the prior rather than independent environmental effects.

  • Primary performance: 76.2 ± 1.7% accuracy was achieved by the bidirectional multimodal framework, an 8.8 percentage-point gain over the EEG-Conformer baseline.Macro F1 was 0.741 ± 0.019 and ROC-AUC was 0.884 ± 0.013.
  • Primary performance: 22.3 ± 1.4% accuracy was achieved by the environment-only baseline, close to chance under label-independent test-time inputs.This indicates little discriminative value for the environmental representation alone in that evaluation condition, without ruling out every form of leakage.
  • Ablations and controls: 71.3% and 70.8% accuracy under random pairing and label-shuffled inputs retained 3.4–3.9 percentage-point gains over EEG-only.These controls indicate that part of the improvement reflected the dual-tower architecture and attention operating on a second input stream.
  • Label-independent diagnostic: 6.3 percentage points separated label-conditioned from label-independent test-time sampling, while 8.8 of the 15.1-point total gap were realized without test-time label access.The label-conditioned condition was a diagnostic upper bound and was not used for the deployed-model result.
  • Environmental domain shift: 72.8% accuracy under the Singapore environmental distribution represented a 3.4 percentage-point shift penalty relative to Astana.The same Astana EEG epochs and labels were evaluated, with the EEG-Conformer frozen and the environmental tower refitted using Singapore data.
  • Environmental attribution: PM2.5 had the highest attribution for anger and sadness, NDVI for relaxed/calm and happiness, and FAR for neutral.GraphLIME attributions describe reliance on the PECM prior, not independent estimates of environmental effects on affect.
  • Spatial sensitivity: 76.2 ± 1.7% accuracy at 1 km2 decreased to 74.0 ± 2.1% at 5 km2 and 72.1 ± 2.3% at 10 km2, while remaining above the 67.4% EEG-only baseline.The reported 4.1-point loss accompanied coarser environmental aggregation.
  • Ablations and controls: 70.9% accuracy under reversed dose-response conditioning remained 3.5 percentage points above EEG-only, while restoring literature-consistent structure recovered additional performance.The authors interpret the residual gain as a limitation of the negative control rather than evidence of a genuine reversed effect.

4 Discussion

The study’s main contribution is a controlled framework for fusing EEG with literature-informed environmental context, while explicitly limiting interpretation because the domains are not individually co-registered. Results support technical feasibility, but expose architecture-, prior-, domain-, temporal-, and privacy-related boundaries.

  • Methodological contribution: The framework integrates spatially structured environmental representations with EEG affective-state classification, improving performance relative to an EEG-only architecture.Its value is methodological: it provides a controlled computational setting for testing neuro-geospatial fusion before jointly collected in-situ data exist.
  • Domain shift: The Singapore substitution reduces accuracy to 72.8%, supporting environmental-domain generalisation as a narrower property than city-agnostic or population-level transferability.The evaluation reuses Astana EEG epochs and labels with a substituted environmental distribution; no independent Singapore EEG data were available.
  • Spatial robustness: 72.1% versus 67.4% accuracy after aggregation to 10 km2 indicates that gains over EEG-only persist at a coarser spatial scale.The result suggests LUR contributes useful spatial priors rather than only fine-scale, scale-dependent detail.
  • Architectural contribution: 3.4–3.9 percentage points of the total 8.8-point gain survives controls disrupting environmental-label structure, while 4.9–5.4 points are attributable to the literature-informed structure.The decomposition indicates that part of the benefit comes from the dual-tower architecture attending to an additional input stream, rather than only from correct environmental content.
  • Interpretive boundary: Environmental context is a computational construction sampled from a city-wide distribution, not an individual’s measured exposure during EEG recording.The benchmark lacks participant-level trajectories, outdoor sensor readings, and timestamped exposure measurements, so the framework is a computational digital twin rather than a test of responses to real outdoor stimuli.
  • Interpretive boundary: The environmental prior is deliberately conditioned on affective labels during training, so performance reflects exploitation of an imposed literature-derived relationship rather than independent discovery.Test-time sampling is label-independent, but the training construction encodes the literature-based association between affective category and environmental regime.
  • Limitations and future work: Static or annually aggregated environmental inputs omit acute exposure-response dynamics, cumulative exposure, lagged effects, and within-person temporal variation.The environmental tower’s GRU currently processes a sequence of length one; future work proposes multi-step sequences from hourly OpenAQ observations.
  • Limitations and future work: Future jointly collected longitudinal EEG-environment datasets could support within-person comparisons while requiring geo-privacy protections for paired location and neural data.Proposed safeguards include spatial aggregation or calibrated spatial noise, differential privacy, clear consent, withdrawal, and deletion procedures.

5 Conclusions

The study combines EEG affective representations with spatially structured environmental context under a leakage-aware protocol for non-co-registered data. It demonstrates technical feasibility, while interpreting gains as model behavior rather than evidence of causal real-world exposure–affect coupling.

  • The protocol combines EEG affective representations with spatially structured environmental context when the domains are not directly co-registered.
  • 58% of the 15.1-point EEG-only-to-upper-bound gap is realised under label-independent inference, while 42% requires privileged test-time information.The corresponding realised and privileged portions are 8.8 and 6.3 points, respectively.
  • 76.2% accuracy is achieved on Astana under label-independent inference.
  • 72.8% accuracy under Singapore environmental-domain replacement is a 3.4-point reduction, interpreted as architectural robustness rather than cross-population generalisation.No independent Singapore EEG data exist.
  • The results demonstrate technical feasibility and generate testable, literature-consistent hypotheses rather than directly measuring real-world brain–environment coupling.The environmental–affect link was constructed from literature dose-response evidence rather than measured in situ.

Declarations

The study is a secondary analysis of the publicly accessible EAV benchmark and uses publicly available environmental data. No new participants were recruited, and the authors report no competing interests.

  • The analysis uses the publicly accessible EAV benchmark and recruited no new human participants.Ethical approval and informed consent for the original data collection were obtained by the dataset creators.
  • OpenAQ, Sentinel-2, Sentinel-5P, and OpenStreetMap provide the other publicly available environmental data used.
  • The manuscript contains no individual person’s data, images, or other identifiable information.
  • Processed environmental grids and spatial features are available through Zenodo, while the third-party EAV dataset is accessed from its original creators under authorised access.
  • The authors declare no competing interests and report support from Vellore Institute of Technology.
  • Author contributions span conceptualisation, methodology, data curation, analysis, supervision, validation, resources, writing, and funding acquisition.

S1 Detailed EEG preprocessing

The EEG preprocessing converts participant recordings into cleaned, non-overlapping five-second epochs and derives differential-entropy and power-spectral-density features across canonical frequency bands.

  • 16,800 EEG epochs are produced from 42 participants by segmenting each trial into four non-overlapping five-second epochs.Each participant contributes 400 epochs; raw EEG has 30 electrodes and is sampled at 500 Hz before preprocessing.
  • Signals are band-pass filtered from 3–50 Hz, downsampled to 100 Hz, and cleaned with Artifact Subspace Reconstruction and FastICA.
  • Differential Entropy and Power Spectral Density are computed across 30 channels and five canonical bands: δ, θ, α, β, and γ.PSD uses Welch’s method with a 256-sample Hanning window and 50% overlap.
  • For band-limited Gaussian EEG, differential entropy is reduced to its Gaussian-form expression.The supplied passage introduces this reduction but does not state the full equation.
  • The neural inputs retain separate DE and PSD tensors with dimensions T × 30 × 5 before concatenating their feature dimensions.

S2 Complete model architecture

The complete architecture uses separate EEG and environmental processing pathways, graph-based spatial encoding, bidirectional cross-modal attention, and a fused classification objective with alignment regularisation.

  • The model is specified as a dual-tower architecture with separate neural and environmental representations before multimodal fusion.
  • The environmental tower combines a six-dimensional static node-feature matrix with a six-dimensional epoch-level PECM sample.
  • The epoch-level PECM sample is broadcast identically across all 100 graph nodes before the first GCN layer.This uniform broadcast carries a population-level environmental prior into a spatially resolved graph without fabricating spatial structure.
  • A two-layer GCN maps node features through dimensions 6 → 64 → 128, followed by a 128-dimensional GRU and global mean pooling over 100 nodes.
  • Bidirectional cross-modal attention uses separate environment-conditioned and EEG-conditioned directions to encode distinct inter-modal interactions.The architecture uses eight attention heads with per-head dimension 32.
  • Classification applies softmax to the fused representation and optimises cross-entropy plus an alignment term based on the squared distance between EEG and environmental embeddings.The concatenation-fusion ablation instead uses a learned projection of the concatenated EEG and environmental vectors.

S3 Hyperparameters

Table S1 lists the complete hyperparameter set used to train the proposed model.

  • Table S1 provides the complete hyperparameter set for the proposed model.

S4 Environmental data processing

Environmental data are calibrated, harmonised, and represented across a common 1 km2 grid, while temporal variability remains largely outside the current static representation.

  • LUR calibration used six Astana and nine Singapore OpenAQ stations, achieving R2 values of 0.74 and 0.68 for Astana PM2.5 and NO2.Singapore calibration achieved R2 = 0.71 for PM2.5 and 0.65 for NO2.
  • Six and nine calibration stations constrain the spatial representativeness of the resulting environmental surfaces.
  • Environmental layers were harmonised onto a common 1 km2 grid by aggregating higher-resolution data.The pipeline combines LUR surfaces, Sentinel-2 NDVI, and OpenStreetMap morphology.
  • All pollution values underlying PECM label-conditional sampling are LUR-estimated.
  • Hourly PM2.5 and NO2 variability shows morning and evening traffic peaks plus an Astana winter PM2.5 elevation absent from Singapore data.These temporal patterns motivate a future multi-step GRU extension but are not propagated into current static node features.

S5 Spatial graph construction

The environmental encoder represents the study area as a 100-cell spatial graph, using six environmental features and distance-weighted spatial connections.

  • The environmental graph contains 100 nodes arranged as a 10×10 grid, with six features per 1 km2 cell.Features include LUR-estimated PM2.5, NDVI, FAR, GVI, LUR-estimated NO2, and O3.
  • Queen’s-contiguity edges connect spatially close cells using inverse-squared-distance weights wij = 1/(d2ij + ϵ).This operationalises Tobler’s first law of geography as a structural inductive bias.
  • A 250 m buffer produced the most stable indicator distributions and was used throughout the main analysis.The 100 m buffer was too narrow, whereas 500 m caused excessive spatial smoothing.

S6 Additional ablation experiments

Additional experiments compare alternative encoders, fusion and alignment objectives, and preprocessing choices to identify which components contribute to performance.

  • Ablations: Removing the GRU changes accuracy by only −0.5 pp relative to the full model.The environmental GRU operates on a static snapshot in the listed configuration.
  • Ablations: Replacing the graph encoder with a flat MLP changes accuracy by −2.3 pp, indicating a contribution from graph structure rather than the temporal module.
  • Fusion and alignment: Replacing bidirectional attention with concatenation fusion changes accuracy by −1.6 pp.The comparison tests fusion design while retaining the main model components.
  • Fusion and alignment: InfoNCE and RMA-projected alternatives change accuracy by −0.8 and −0.2 pp, respectively.The results indicate that the cross-modal objective and attention mechanism matter more than the specific closely performing alignment variant.
  • Controls: The reversed dose-response control inverts the quartile column while leaving every other pipeline element unchanged.Upper-quartile environmental values are computed separately for each city, while neutral epochs use interquartile-range values.
  • Alternative objectives: RMA alignment did not yield a consistent additional gain over the deployed squared-ℓ2 objective.Canonical-correlation-style and SPD-aware alternatives likewise did not outperform it by a meaningful margin.

S7 Latent-space visualisation

UMAP visualisations show how the cross-modal alignment objective changes the learned joint representation geometry, bringing neural and environmental embeddings into closer, more separable affective-state clusters. This visualisation is descriptive and does not independently confirm an environmental mechanism.

  • UMAP visualises representations learned by the model rather than performing the alignment itself.
  • Before training, neural and environmental embeddings occupy distinct, poorly overlapping regions of the projected space.
  • After alignment, the five affective-state clusters become more compact and separable, with neural and environmental centroids visibly closer together.
  • The observed overlap reflects learned representation geometry, not independent evidence of an environmental mechanism.

S8 Additional spatial sensitivity analysis (MAUP)

The spatial sensitivity analysis tests whether multimodal performance depends on fine environmental spatial resolution. Accuracy declines as resolution coarsens, but the multimodal model remains above the EEG-only baseline at 10 km2.

  • PM2.5 and NO2 layers were aggregated from native 1 km2 LUR resolution to 5 km2 and 10 km2 using block averaging.
  • Performance degrades roughly monotonically as spatial resolution coarsens, with a total loss of 4.1 percentage points from 1 km2 to 10 km2.
  • 72.1% accuracy is maintained at the coarsest scale, above the EEG-only baseline of 67.4%.
  • NDVI, FAR, and GVI were held at native resolution during the spatial scale sensitivity analysis.
Loading 2608.20807v1…