Source-linked AI summary
Score-Based Generative Data Assimilation for Integrating Aggregated Surveillance Data into Agent-Based Models in Epidemic Tracking
Siming Liang, Jacob Hauck, Minglei Yang, Adam Spannaus, Heidi Hanson, Guannan Zhang
TL;DR
The paper targets the difficulty of inferring epidemic burden and transmission heterogeneity from noisy aggregated surveillance while calibrating partially observed ABMs. GenDA jointly updates macrostates and heterogeneous parameters, then reconciles those corrections with agent-level configurations. In synthetic controlled and geographically explicit experiments, it recovers burden trends and hotspot structures while improving forecasts after assimilation compared with state-only assimilation.
Problem
Aggregated surveillance obscures latent epidemic patterns, while ABM contact and behavior heterogeneity makes online joint state–parameter calibration difficult.
Method
GenDA combines training-free score-based macrostate filtering, discrepancy-informed direct parameter updates, and macro–micro reassignment for partially observed epidemic ABMs.
Results
GenDA recovers burden trends and dominant hotspot structures in controlled and geographically explicit synthetic ABMs while improving post-assimilation forecasts over state-only assimilation.
Takeaways & Limitations
The framework supports coherent aggregate-level calibration of partially observed epidemic ABMs without requiring exact recovery of individual trajectories.
Takeaways & Limitations
Evidence is synthetic; empirical validation and larger-scale scalability remain unestablished, and parameter identifiability can weaken under sparse supports or model-form error.
Abstract
from arXiv · showhide
Reliable epidemic monitoring often requires inferring regional infection burden and transmission heterogeneity from noisy, spatially aggregated, and potentially sparse surveillance data. Agent-based models (ABMs) are attractive for this task because they represent individual behavior, contact heterogeneity, and localized interventions, but these same features make them difficult to calibrate online. We develop a generative AI-based data-assimilation (GenDA) framework for partially observed epidemic ABMs that estimates both the epidemic state and a heterogeneous parameter field while respecting the gap between observable macrostates and latent agent-level microstates. GenDA combines a training-free, score-based generative update for macrostate correction with a direct parameter update based on macrostate discrepancies, followed by a macro-micro reassignment step that restores consistency with the ABM. In controlled and geographically explicit synthetic experiments, the framework recovers regional epidemic burden, dominant hotspot structures, and effective transmission heterogeneity from aggregated observations, while improving post-assimilation forecasts relative to state-only assimilation.
1. Introduction.
The paper addresses sequential calibration of partially observed epidemic ABMs, where aggregated surveillance must inform both regional states and heterogeneous parameters despite latent agent-level dynamics. GenDA combines non-Gaussian macrostate filtering, discrepancy-informed parameter updates, and macro–micro reassignment to extract epidemiologically meaningful information.
- ABMs represent individual contacts, movement, behavioral adaptation, and localized interventions, but these features complicate online calibration.
- Aggregated surveillance creates a coupled inference problem linking coarse regional macrostates, latent agent configurations, and spatially heterogeneous parameters.
- GenDA uses a non-Gaussian score-based macrostate filter and transfers forecast–analysis discrepancies into parameter-space pseudo-observations without direct parameter observations.
- A macro–micro reassignment step restores consistency between the assimilated aggregate state and agent-level simulator configurations.
- The framework is validated in controlled and geographically explicit ABMs, including comparison with score filtering that updates the epidemic state only.
2. Problem setting.
The framework uses a stochastic, spatially heterogeneous ABM to forecast epidemics and assimilates surveillance data reported at coarser regional resolutions. Its central task is to update aggregated epidemic states and heterogeneous parameters, then reconcile those corrections with agent-level states.
- Conceptual reference: A spatially heterogeneous SIR model provides an interpretable conceptual reference but is not used directly for simulation or inference.The reference describes transmission, recovery, and spatial spread through aggregate quantities and spatially varying coefficients.
- Agent-based forecast model: The forecast model is a stochastic ABM representing heterogeneous contacts, mobility, agent attributes, and geographically localized interventions.Agent-specific parameters may be shared across spatial, demographic, behavioral, or other groups to reduce inferential dimension.
- Assimilation objective: The inferential task updates random regional macrostates and the parameter field from aggregated observations, then modifies agent states to match the corrected regional counts.This reconciliation is difficult because both the simulated agent state and aggregate statistics are stochastic.
- Data-assimilation motivation: Uncertain and time-varying parameters, incomplete behavioral and mobility information, model-form error, and intrinsic stochasticity can make simulated trajectories drift from the real epidemic.This drift motivates sequential calibration and integration of surveillance data.
- Macrostate observations: The central resolution mismatch is between agent-level ABM states and surveillance statistics reported over disjoint spatial sub-regions.Regional macrostates aggregate agent information, while observations may select compartments or apply reporting transformations and include observation error.
3. Score-based generative data assimilation for joint state and parameter estimation.
GenDA jointly updates aggregated epidemic macrostates and heterogeneous ABM parameters through score-based filtering, then restores consistency with latent agent states. Its assimilation cycle combines nonlinear macrostate correction, discrepancy-informed parameter inference, and local macro–micro reassignment.
- Framework overview: Each assimilation cycle filters regional macrostates, infers the ABM parameter field, and modifies agent disease states to enforce macro–micro consistency.The framework is designed for joint estimation using aggregated observations.
- Macrostate estimation: Score-based nonlinear filtering combines an ensemble prior score with the surveillance likelihood gradient during reverse diffusion to sample the posterior macrostate ensemble.Forward diffusion transports forecast samples to a standard Gaussian reference, while reverse diffusion returns posterior samples.
- Parameter estimation: Exploratory parameter realizations are weighted by forecast agreement with the updated macrostate, producing a parameter pseudo-observation for score-based parameter filtering.The exploration spread is controlled by γt and may decrease over time from coarse exploration to refined estimation.
- Macro–micro consistency: After assimilation, agent disease states are locally reassigned to match member-specific updated macrostates, so subsequent simulations start from exact micro–macro consistency.Updated parameters are assigned to corresponding agents or parameter-sharing groups.
4. Numerical experiments: recovering epidemic burden, hotspot structure, and transmission heterogeneity.
Across controlled and geographically explicit synthetic experiments, GenDA improved epidemic state tracking and online parameter calibration from aggregated observations. Joint state–parameter assimilation preserved more accurate post-assimilation forecasts and spatial patterns than state-only assimilation, while long-horizon local errors remained sensitive to stochastic agent dynamics.
- Experimental design: The experiments evaluated state estimation and online parameter estimation in stochastic epidemic ABMs using Mesa implementations.The study included a controlled uniform-grid setting and a geographically explicit ABM-GEO setting.
- Controlled validation: The controlled setup used 4,000 agents on a 20 × 20 grid, 100 aggregated 2 × 2 observation blocks, and 200 spatially varying transmission and recovery parameters.Each observation block had independently sampled transmission and recovery rates, creating a heterogeneous inference problem.
- Controlled validation: State-only assimilation tracked the reference during the first 30 time steps but deteriorated rapidly after assimilation stopped because parameters remained misspecified.This contrasts with the no-assimilation baseline, whose ensemble mean substantially deviated from the reference trajectory.
- Controlled validation: Joint state–parameter assimilation maintained accurate prediction after the assimilation window and reduced parameter RMSE, indicating online calibration of heterogeneous parameters.Joint filtering also narrowed macrostate ensemble spread as observations constrained both states and inferred parameters.
- Block-level spatial accuracy and uncertainty: At T = 30, state-only and joint assimilation both achieved low spatial errors, but joint assimilation retained lower errors at T = 50 and T = 100.Long-horizon local spatial errors still grew because stochastic mobility, contacts, and discrete interactions affect spatial allocation.
- Geographically explicit experiments: In the geographically explicit ABM-GEO experiment, joint assimilation reduced parameter error and produced more stable aggregate epidemic estimates under sparse, uneven observations.At T = 30, both assimilative cases recovered dominant infected-density patterns, while joint updating better preserved broad high-risk regions and elongated hotspot structures during prediction.
- Interpretation and scope: The framework targets recoverable macro-scale epidemic structure and representative spatial realizations rather than exact reconstruction of individual agent trajectories.This scope reflects the use of sparse, noisy, aggregated observations and a stochastic agent-based simulator.
5. Discussion and Conclusion.
The experiments support aggregate-level calibration of partially observed epidemic ABMs, with joint state–parameter filtering preserving regional patterns and improving forecasts. Evidence remains synthetic and moderate-scale, with empirical validation and broader scalability still needed.
- Results: Case 3 maintains smaller post-assimilation errors and dominant regional patterns than state-only Case 2 in ABM-GEO predictions.The comparison covers infected-density predictions at T = 40, 50, and 60 after assimilation ends at T = 30.
- Results: Joint state–parameter filtering preserves burden trends and dominant hotspot structures after the assimilation window.The pattern persists under uneven population density, spatially uneven observation support, stochastic mobility, and hub-mediated contact structure.
- Interpretation: The framework should be evaluated using support-level burdens, spatial risk gradients, and forecast distributions rather than exact reproduction of one realized agent trajectory.Inferred transmission and recovery fields are effective parameters summarizing heterogeneous processes at the chosen sharing scale.
- Limitations: Both experiments are synthetic, and the Toronto case uses real geometry but not real population, facility, or surveillance data.Empirical validation remains necessary before operational use.
- Limitations: Parameter identifiability may weaken when distinct parameter fields produce similar aggregate trajectories, especially under sparse supports, model-form error, or misspecified observation noise.The exact local reassignment also assumes disjoint supports; overlapping supports require a joint constrained allocation.
- Future work: A primary next step is integrating GenDA with the ENABLE high-performance agent-based framework and validating estimates against empirical surveillance.Further work should also evaluate support geometry, reporting error, and overlapping constraints.
- Conclusion: GenDA provides a testable foundation for larger simulators and empirical surveillance data by combining non-Gaussian macrostate assimilation, direct parameter updates, and macro–micro consistency.The stated scope is extracting useful epidemic summaries where exact fine-scale predictability is fundamentally limited.
Appendix A. A Reader’s Guide for Life-Science Audiences.
The ABM evolves individual agents while the available data provide regional surveillance summaries, so the method infers macrostates rather than exact hidden agent histories.
- Interpretation: The method estimates sub-region epidemic counts and a heterogeneous parameter field instead of reconstructing every agent’s hidden history.Many agent-level configurations can produce the same regional counts in a stochastic epidemic model.
A.1. Interpreting the filtering cycle.
The reader guide directs audiences to foundational resources on epidemic ABMs, large-scale modeling platforms, and uncertainty quantification for stochastic ABM calibration.
- Further reading: The guide recommends general ABM texts, epidemic-ABM reviews, and platforms including EpiPredict, UVA-EpiHiper, and ENABLE.It also points readers toward references on uncertainty quantification and calibration for stochastic ABMs.