Source-linked AI summary
Simulation of an SEIR infectious disease model on the dynamic contact network of conference attendees
Juliette Stehlé, Nicolas Voirin, Alain Barrat, Ciro Cattuto, Vittoria Colizza, Lorenzo Isella, Corinne Régis, Jean-François Pinton, Nagham Khanafer, Wouter Van den Broeck, Philippe Vanhems
TL;DR
The paper asks how much temporal and duration detail contact data must retain to model infectious-disease spread realistically. It uses RFID-based conference interactions to compare a full dynamic network with daily aggregated alternatives in SEIR simulations. Retaining daily contact durations approximates the full network, whereas topology-only homogeneous aggregation produces larger and faster outbreaks.
Problem
Few empirical contact studies provide objective, person-to-person measurements with sufficient temporal detail, leaving the required modeling resolution uncertain.
Method
The study maps high-resolution conference face-to-face interactions into a dynamic network and compares SEIR simulations with two daily aggregated network representations.
Results
Retaining heterogeneity in cumulated contact duration gives crucial information, while HOM differs from HET and DYN in simulated spreading outcomes.
Takeaways & Limitations
Daily aggregation that preserves contact duration can approximate the full-resolution network on the timescales considered, whereas topology-only representation does not reproduce epidemic size.
Takeaways & Limitations
The study’s discussion identifies caveats associated with the data and its interpretation.
Abstract
from arXiv · showhide
The spread of infectious diseases crucially depends on the pattern of contacts among individuals. Knowledge of these patterns is thus essential to inform models and computational efforts. Few empirical studies are however available that provide estimates of the number and duration of contacts among social groups. Moreover, their space and time resolution are limited, so that data is not explicit at the person-to-person level, and the dynamical aspect of the contacts is disregarded. Here, we want to assess the role of data-driven dynamic contact patterns among individuals, and in particular of their temporal aspects, in shaping the spread of a simulated epidemic in the population. We consider high resolution data of face-to-face interactions between the attendees of a conference, obtained from the deployment of an infrastructure based on Radio Frequency Identification (RFID) devices that assess mutual face-to-face proximity. The spread of epidemics along these interactions is simulated through an SEIR model, using both the dynamical network of contacts defined by the collected data, and two aggregated versions of such network, in order to assess the role of the data temporal aspects. We show that, on the timescales considered, an aggregated network taking into account the daily duration of contacts is a good approximation to the full resolution network, whereas a homogeneous representation which retains only the topology of the contact network fails in reproducing the size of the epidemic. These results have important implications in understanding the level of detail needed to correctly inform computational models for the study and management of real epidemics.
Background
The paper addresses how much detail contact-network data must retain to support realistic epidemic simulations. It focuses on temporal contact heterogeneity and constraints, using high-resolution conference interactions to compare network representations.
- Modeling requirements: Contact networks for epidemic modeling should represent variation in contact duration and frequency, as well as causal ordering between interactions.These features determine which transmission sequences are possible in a dynamic contact network.
- Motivation: Empirical contact data often lack person-to-person, temporal, and longitudinal detail, while self-reported methods can involve bias and limited representativeness.Diary-based approaches commonly cover few people and limited time snapshots, making short respiratory-route encounters difficult to report exhaustively.
- Research question: The paper asks what level of contact-data detail is needed for realistic epidemic simulations without making models unnecessarily opaque.The study contrasts overly coarse homogeneous mixing with extremely detailed representations that can obscure the effects of individual assumptions.
- Study aim: Using high-resolution face-to-face conference data, the study evaluates temporal aspects, heterogeneities, and constraints in simulated infectious-disease transmission.RFID-derived interactions are mapped to a dynamic contact network and compared with two daily aggregated projections.
- Study aim: The simulations are intended to identify the contact-data detail needed to inform computational models for public-health problems adequately and realistically.The analysis compares simulated outbreaks across explicit dynamic and aggregated contact representations.
Methods
The study builds a dynamic contact network from RFID-recorded face-to-face interactions and compares it with daily networks that retain different amounts of temporal information. It extends the observations over longer epidemic timescales before simulating SEIR spread.
- Network representations: DYN preserves each contact’s participants, start and end times, chronology, duration, and causality constraints at 20-second resolution.This representation can exclude transmission paths that an aggregated static network would allow.
- Network representations: HET aggregates contacts by day while retaining who met whom and each pair’s total face-to-face duration, but discards contact order.The two-day conference produces one HET network per day.
- Network representations: HOM uses the same daily interaction topology as HET but assigns every link the average contact duration.Thus HET and HOM differ only in their link weights.
- Longer-timescale extension: The two-day contact data are extended longitudinally by repeating sequences or randomly reshuffling participant identities before constructing DYN, HET, and HOM networks.Repetition preserves individual sequences but becomes unrealistic over time, whereas RAND-SH removes correlations between successive sequences.
Results
The empirical conference network was heterogeneous in topology, contact duration, and individual activity. In SEIR simulations, HOM produced faster, larger outbreaks, whereas HET closely matched DYN despite discarding contact order.
- Empirical network: 28,540 face-to-face contacts among 405 attendees formed a clustered small-world network with average clustering 0.28 and average shortest path 2.2.The clustering coefficient exceeded 0.07 for a same-size random network with the same average degree.
- Empirical network: Daily individual activity was correlated across days, with Pearson coefficients of 0.37 for degree and 0.52 for total interaction time; repeated contacts accounted for 12%.The repeated-contact fraction was independent of degree.
- Epidemic simulations: HOM had lower extinction probability and higher final outbreak size than HET and DYN, while HET and DYN produced similar outcomes.HOM spreading also evolved slightly faster, whereas HET and DYN had very similar temporal characteristics.
- Epidemic simulations: Peak-time differences were small, so even the least-informative network estimated peak time well relative to the full-information contact network.HOM generally reached the epidemic peak first.
- Extension procedures: RAND-SH spreading was slightly slower but lasted longer, with spreading efficiency increasing as tag identities were shuffled more extensively.The extension procedure changes longitudinal correlations while preserving the overall contact sequence.
Discussion
The simulations compare dynamic, heterogeneous, and homogeneous contact representations using high-resolution conference interaction data. Daily cumulative contact duration largely reproduces epidemic propagation, whereas ignoring duration heterogeneity substantially alters outbreak size; precise contact ordering has little effect on the studied timescales.
- Data and network representations: 405 conference volunteers supplied high-resolution face-to-face interaction measurements for simulations on dynamic and aggregated contact networks.The data collection covered a two-day conference and captured contacts with fine temporal and spatial resolution.
- Epidemic comparison: 36%–47% of simulations ended in extinction, while 34%–49% became large outbreaks with 51%–80% attack rates across all three networks.These outcomes were described as frequent and explosive, respectively, and were consistent with previous work.
- Epidemic comparison: The homogeneous network produced systematically more infected individuals than the heterogeneous and dynamic networks.HOM omits contact-duration heterogeneity and the dynamical aspect of contacts.
- Epidemic comparison: Contact-duration heterogeneity was associated with lower spread because unequal time allocation across contacts reduces available spreading routes.Disregarding duration heterogeneity can strongly change estimated case counts, making daily cumulative contact time important for modeling.
- Epidemic comparison: The peak time changed only slightly in the homogeneous network, indicating that limited contact information can still estimate epidemic timescales well.Thus, epidemic size is more sensitive than peak timing to omitted contact-duration heterogeneity in these simulations.
- Temporal resolution: Daily cumulative contact time was sufficient to reproduce propagation patterns on the studied timescales, while precise contact order was not generally needed.The conclusion applies to spreading processes on the order of days, not extremely fast processes, for which a crossover is expected.
- Temporal resolution: Repeated contacts accelerated early spread, whereas new contacts between successive days ultimately produced a larger attack rate.These differences show that the fractions of repeated and new contacts matter when extending the two-day dataset.
- Limitations: The study is constrained by two-day sampling, incomplete RFID coverage, nighttime isolation assumptions, and partial participation among conference attendees.Unmonitored contacts and unsampled attendees may underestimate spreading opportunities, while assumed isolation during 56% of the time may increase extinction probability.
Conclusions
The study emphasizes that contact-duration heterogeneity matters for epidemic modeling, while minute-level contact ordering is less essential on multi-day or multi-week timescales.
- Small differences between HET and DYN spreading indicate that minute-level contact ordering is not essential over several days or weeks.The conclusion contrasts detailed temporal ordering with the longer epidemic timescale considered.
- Strong differences between HOM and the other networks underline the need to include contact-duration heterogeneity rather than assume homogeneous contacts.The comparison identifies duration heterogeneity as important for modeling disease spread.
- The combined comparisons of HET, DYN, and HOM networks and data-extension procedures assess the contact-pattern detail needed for epidemic-spreading models.The study also examines how different procedures for extending the data affect spreading dynamics.
- RFID-based data collection provides access to information needed for modeling and permits simulation of very fast spreading processes.The infrastructure is presented as effective for capturing relevant contact information and fast dynamics.
- The conclusions are bounded by the possibility that extremely fast spreading processes enter a different regime, with a crossover left for future investigation.The paper also calls for longer measurements and data from workplaces, schools, hospitals, and other settings.
- More experiments are needed to understand how time-limited datasets can be extended into realistic datasets across samples and locations.The authors connect this need to anticipating preventive-measure impacts and informing control strategies.
Authors' contributions
The authors divided the work across experiment design, data collection, data analysis, and paper writing, with all authors approving the final manuscript.
- The experiments were conceived and designed by JS, NV, AB, CC, VC, LI, CR, JFP, WVdB, and PV.
- Data collection was performed by NV, AB, CC, CR, JFP, NK, WVdB, and PV.
- Data analysis was conducted by JS, NV, AB, CC, VC, LI, and JFP.
- The paper was written by JS, NV, AB, CC, LI, JFP, and PV, and all authors approved the final manuscript.
Dynamics of Person-to-Person Interactions from Distributed RFID
This section cites prior work on evolving social networks, face-to-face behavioral networks, high-resolution human contacts, and dynamical processes on complex networks.
- Cited work addresses empirical analysis of an evolving social network.
- The references include work on mobile communication-network structure and tie strengths.
- Related studies examine face-to-face behavioral networks and high-resolution human contact networks for infectious-disease transmission.
- Other cited research concerns computational social science, dynamical processes on complex networks, and epidemic reproduction numbers.
- The bibliography also includes studies of epidemic models, contact repetition and clustering, and prioritizing healthcare-worker vaccinations using social-network analysis.
Figures
The figures characterize contact durations, R0, epidemic size, peak timing, and temporal spreading across network types and parameter scenarios. Reported summaries include a 49-second average contact duration and a 112-second standard deviation.
- 49 seconds is the average duration of contacts between individuals.
- 112 seconds is the reported standard deviation of contact duration.
- Figures 2 and 3 report R0 distributions for the HOM, HET, and DYN networks under the stated scenarios.Figure 2 concerns the REP procedure; Figure 3 presents boxplots across scenarios and network types.
- Figures 4 and 5 examine final case counts and prevalence peak times across the three network types and scenarios.The peak-time analysis includes only runs with AR>10%.
Tables
Table 1 summarizes final case-number distributions for three network types across four scenarios using 5000 runs. The listed outcome categories are 1 to 10, 11 to 40, and more than 40 final cases.
- Table 1 reports the distribution of final cases for three network types across four scenarios.The table is based on 5000 runs with a dynamic contact network.
- The accompanying network comprises 405 participating attendees, and 90%-CI denotes a 90% confidence interval.
- The final-case categories are 1 to 10, 11 to 40, and more than 40 cases.
Additional Material
The additional material details how contact data are reshuffled while preserving empirical day-to-day repetition, and describes the resulting contact graphs and epidemic-summary statistics.
- Data extension procedure: CONSTR-SH reshuffles tag identities to generate artificial contact data while preserving realistic correlations between successive daily networks.The procedure compares repeated-contact fractions with the empirical value and accepts exchanges probabilistically to reduce their squared deviation.
- Data extension procedure: femp measures the average fraction of an individual’s day-2 contacts that were also present among that individual’s day-1 contacts.femp = 0 indicates entirely new contacts on day 2, whereas femp = 1 indicates identical contact sets across both days.
- Data extension procedure: The reshuffling algorithm repeatedly exchanges randomly selected tag identities, computes f and (f – femp)2, and accepts exchanges with probability controlled by b.Increasing b gradually produces reshufflings with very low deviation from the empirical repeated-contact fraction.