Source-linked AI summary
Can co-location be used as a proxy for face-to-face contacts?
Mathieu Génois, Alain Barrat
TL;DR
The paper examines whether coarse, temporally resolved co-presence data can substitute for detailed face-to-face contact data in network analysis and epidemic modeling. It compares both network types across contexts and evaluates down-sampled surrogate contacts, finding partial recovery of contact features but strong context dependence in epidemic and containment results.
Problem
It is unclear how much co-presence data can reveal about actual face-to-face contacts, which matter for studying population structure and modeling information or epidemic spreading.
Method
The authors compare temporally resolved face-to-face and coarse co-presence networks from several SocioPatterns contexts and test sampling methods that create surrogate contact networks.
Results
Sampling improves agreement with real contacts for several network statistics and epidemic simulations, but performance varies by context and central-node identification remains limited.
Takeaways & Limitations
Co-presence can approximately recover some aggregate contact properties for data-driven models, but it is not a reliable systematic substitute for detailed contacts or containment-strategy evaluation.
Takeaways & Limitations
The conclusions are context-dependent, and co-presence-based surrogates do not reliably reproduce detailed contacts or identify the most central nodes.
Abstract
from arXiv · showhide
Technological advances have led to a strong increase in the number of data collection efforts aimed at measuring co-presence of individuals at different spatial resolutions. It is however unclear how much co-presence data can inform us on actual face-to-face contacts, of particular interest to study the structure of a population in social groups or for use in data-driven models of information or epidemic spreading processes. Here, we address this issue by leveraging data sets containing high resolution face-to-face contacts as well as a coarser spatial localisation of individuals, both temporally resolved, in various contexts. The co-presence and the face-to-face contact temporal networks share a number of structural and statistical features, but the former is (by definition) much denser than the latter. We thus consider several down-sampling methods that generate surrogate contact networks from the co-presence signal and compare them with the real face-to-face data. We show that these surrogate networks reproduce some features of the real data but are only partially able to identify the most central nodes of the face-to-face network. We then address the issue of using such down-sampled co-presence data in data-driven simulations of epidemic processes, and in identifying efficient containment strategies. We show that the performance of the various sampling methods strongly varies depending on context. We discuss the consequences of our results with respect to data collection strategies and methodologies.
1 Introduction
The paper asks whether coarse co-location data can recover population-level face-to-face contact patterns despite limits on collecting high-resolution contact data. It compares temporal co-presence and face-to-face networks, then evaluates down-sampling methods for constructing surrogate contacts.
- High-resolution contact measurement is difficult at arbitrarily large population sizes, motivating contact proxies from lower-resolution or alternative data sources.
- The study targets overall population contact patterns rather than inferring specific contacts between individual pairs.
- The authors use temporally resolved SocioPatterns data containing both face-to-face encounters and coarse individual location tracking across multiple contexts.
- Co-presence and face-to-face temporal networks share important structural and statistical properties, but co-presence is much denser because it uses lower spatial resolution.
- Several down-sampling methods are evaluated for transforming the denser co-presence signal into surrogate contact networks.
2 The co-presence network
The study constructs temporally resolved face-to-face contact and co-presence networks from wearable-sensor and RFID data across several social contexts. Co-presence broadly resembles contact patterns but produces denser networks and can overestimate contact durations.
- Data sets and network construction: Face-to-face proximity is detected at 1.5 m, while co-presence requires individuals to share the same exact set of RFID readers.Readers cover approximately 30 m in open space, with building reception ranges depending on structure and materials.
- Data sets and network construction: The data include workplaces, a hospital, a primary school, a conference, and a high school, with contact and co-presence networks resolved in 20-second windows.Workplace data were collected in two different years, InVS13 and InVS15.
- Temporal and aggregated network properties: Co-presence and contact events have similarly shaped broad temporal distributions, providing approximate information about contact-signal functional forms.The comparison covers event and inter-event durations, cumulative event durations, and numbers of events for individual pairs.
- Temporal and aggregated network properties: Co-presence distributions are typically broader with heavier tails, and co-presence-only data overestimate contact and aggregate durations.Differences are strongest in the primary-school data, where one reader covers the schoolyard and some readers cover multiple classrooms.
- Temporal and aggregated network properties: Co-presence daily networks are much denser than contact networks, with larger average degree, clustering coefficient, and cliques.In some school and conference cases, aggregated co-presence networks are close to fully connected.
- Temporal and aggregated network properties: Despite density differences, contact and co-presence matrices show similar group structures, especially for hospital data and for classes and class groups in the high school.The matrices compare average link densities between and within groups across days.
- Temporal and aggregated network properties: Except in the office cases, median contact activity correlates with the number of individuals present and follows a power-law shape with an exponent around 1.5.The relationship has huge context-dependent fluctuations, including many zero-contact instances at potentially large occupancy in InVS15.
3 Sampling co-presence data
The paper down-samples coarse co-presence data into surrogate contact networks using three methods and compares their temporal, structural, and centrality properties with real face-to-face contacts. Surrogates recover some network features, but accuracy varies by method and context, especially for node centrality.
- Sampling methods: Three methods sample individual co-presence times, complete sampled pair trajectories, or sample whole co-presence events to create surrogate contact networks.Each method uses co-presence data differently, ranging from isolated time points to complete events with durations.
- Temporal properties: Contact activity timelines are broadly recovered, but detailed intra-day variations are not always reconstructed, with strongest deviations for method 2 in conference and high-school data.The primary-school case is an exception for overall timeline recovery.
- Temporal properties: Method 1 produces exponential surrogate contact durations, whereas methods 2 and 3 can preserve broader duration distributions depending on context.Method 3 matches real durations for InVS13, LH10, and SFHH but resembles method 2 in other cases.
- Network structure: Surrogate networks reproduce daily contact-matrix similarity and aggregated weights better than raw co-presence data, despite context-dependent degree over- or underestimation.Method 1 overestimates degrees, method 2 generally shifts them lower, and method 3 varies by context.
- Network structure: Across increasing aggregation windows, sampled networks remain closer to contact data than co-presence networks, with strength usually recovered better than degree.The best method for degree evolution remains context dependent.
- Node centralities: No sampling method reliably identifies the most central nodes at low N; at the top 50%, similarity is often only about 0.5, identifying roughly 25% of the most central nodes.Results generally exceed random baselines but do not outperform rankings from the full co-presence network.
4 Using surrogate contact data in epidemic simulations
The paper tests surrogate contact networks in SIR epidemic simulations and evaluates whether they reproduce epidemic risk and vaccination-strategy rankings. Results are uneven and context dependent: some methods estimate outbreak risk or preserve strategy rankings in particular settings, but none is uniformly reliable.
- Limitations: None of the sampling methods describes all relevant contact-network features accurately, combining useful similarities with potentially important discrepancies.This limitation motivates cautious use of surrogate data in spreading simulations and containment evaluation.
- Epidemic simulations: The study evaluates surrogate networks with SIR simulations and measures large-outbreak probabilities and outbreak-size distributions.Simulations vary the reproductive number and compare empirical contacts with 100 surrogate-network instances per sampling method.
- Epidemic simulations: Surrogate epidemic-risk estimates depend strongly on both sampling method and context: method 1 often overestimates risk, while methods 2 and 3 show mixed over- and underestimation.Method 2 performs well for conference, school, and high-school data but underestimates risk for offices and hospitals; method 3 is accurate for offices and hospitals but not uniformly elsewhere.
- Containment strategies: Vaccination strategies are compared by outbreak-probability and median-outbreak-size ratios, with rankings ordered by their efficiency in the real contact network.The evaluation includes random, centrality-based, and group-based vaccination strategies.
- Containment strategies: Strategy rankings can be preserved in some contexts but strongly reshuffled in others; for SFHH, method 1 reaches Kendall’s tau = 0.818 for outbreak-size ranking.Thiers13 provides an example where the ranking is strongly reshuffled.
5 Discussion and conclusion
Down-sampled co-presence can approximate some aggregate properties of face-to-face contact networks and support epidemic simulations in selected contexts, but it does not reliably reproduce detailed contacts or containment-strategy outcomes.
- Contact-network properties: Surrogate networks produced by down-sampling co-presence data generally match real contact-duration statistics more closely than raw co-presence networks.Aggregate contact durations, which influence epidemic processes, can be approximately recovered through simple sampling.
- Contact-network properties: Several structural properties, including average degree, average clustering, and largest clique and core sizes, vary strongly by context.The accuracy of these properties is not consistent across the studied settings.
- Contact-network properties: The most central nodes in the real contact network are not identified better than with bare co-presence information.Down-sampling therefore provides no consistent improvement for central-node identification.
- Spreading-process simulations: Epidemic simulations on one sampling method can approximate real-data results, whereas other methods overestimate or underestimate them, depending on context.All sampled methods perform closer to the real contact network than raw co-presence, which strongly overestimates contacts and epidemic risk.
- Implications: Co-presence data is not reliably or systematically sufficient to reproduce detailed contacts or predict spreading and containment outcomes for short-distance contagion.The authors recommend collecting both high-resolution relative-position data and coarser co-presence information.