Source-linked AI summary

Temporal motifs reveal homophily, gender-specific patterns and group talk in mobile communication networks

Lauri Kovanen, Kimmo Kaski, János Kertész, Jari Saramäki

arXiv:1302.2563v1physics.soc-phcs.SIphysics.data-an

TL;DR

The paper asks whether multi-person interaction patterns contain temporal structure that aggregate networks based only on interaction frequencies cannot capture. It applies colored temporal motifs and a network-conditioned null model to mobile-phone records, finding node-type effects, temporal homophily, gender differences, and dense-versus-sparse temporal variation. The framework is constrained by current applicability to data with at most one event per node at a time or events without duration.

  • Problem

    The paper addresses whether interaction sequences involving multiple individuals reveal structure beyond aggregate-network interaction frequencies, including temporal homophily beyond aggregate homophily.

  • Method

    The study applies colored temporal motifs to mobile-phone records and develops a null model conditioned on the weighted aggregate network to test node-type effects.

  • Results

    About 35% of motifs have |z| > 1.96, and the analysis finds node-type effects, temporal homophily, gender differences, and distinct patterns in dense and sparse network regions.

  • Takeaways & Limitations

    Temporal differences identified by the framework are independent of the aggregate network and complement information from the weighted aggregate network.

  • Takeaways & Limitations

    The framework currently applies only to data where nodes have at most one event at a time or events have no duration.

Abstract

from arXiv · show

Electronic communication records provide detailed information about temporal aspects of human interaction. Previous studies have shown that individuals' communication patterns have complex temporal structure, and that this structure has system-wide effects. In this paper we use mobile phone records to show that interaction patterns involving multiple individuals have non-trivial temporal structure that cannot be deduced from a network presentation where only interaction frequencies are taken into account. We apply a recently introduced method, temporal motifs, to identify interaction patterns in a temporal network where nodes have additional attributes such as gender and age. We then develop a null model that allows identifying differences between various types of nodes so that these differences are independent of the network based on interaction frequencies. We find gender-related differences in communication patters, and show the existence of temporal homophily, the tendency of similar individuals to participate in interaction patterns beyond what would be expected on the basis of the network structure alone. We also show that temporal patterns differ between dense and sparse parts of the network. Because this result is independent of edge weights, it can be considered as an extension of Granovetter's hypothesis to temporal networks.

I. TEMPORAL MOTIFS IN COLORED NETWORKS

Colored temporal motifs classify connected interaction patterns by both event order and node attributes. The framework uses time-window connectivity and valid temporal subgraphs to count these motif classes.

  • Definitions: A temporal network represents interactions as time-stamped, duration-bearing events between nodes, which may be assigned colors representing node types.Each event is directed from one node to another and includes a start time and duration.
  • Connectivity: Events are Δt-adjacent when they share a node and occur within the chosen time window.Sequences of such adjacencies define Δt-connected event sets.
  • Connectivity: Valid temporal subgraphs are connected event sets whose events are consecutive for each node.This validity condition restricts the event sets used for motif construction.
  • Motif classes: A temporal motif is an equivalence class of valid temporal subgraphs with isomorphic colored graphs and identical event order.Motif count C(m) is the number of valid temporal subgraphs in the class.

II. NULL MODEL FOR DIFFERENCES BETWEEN NODE TYPES

The null model provides reference motif counts under the assumption that node types do not affect motifs, while preserving the weighted aggregate network. This conditioning isolates type-related differences from activity and connectivity structure.

  • Motivation: Motif counts require comparison with a reference because an isolated count does not indicate whether a motif is unusually common or rare.The null model supplies that comparison.
  • Null model: The null model estimates motif counts under no dependence on node types while conditioning on the weighted aggregate network.This preserves the observed aggregate-network structure when generating reference counts.
  • Null model: Conditioning on the aggregate network prevents observed differences from being explained by node-type frequencies, activity levels, or preferred connectivity patterns.Differences between observed and reference counts therefore target effects beyond those aggregate-network properties.

A. Synthetic data

Synthetic data tests whether the null model detects known type-dependent temporal patterns. The model identifies overrepresented causal motifs even when the weighted aggregate network itself has no distinguishing structure.

  • Synthetic setup: Synthetic data contains red and blue nodes whose causal chains are more common when the first event connects same-color nodes.The weighted aggregate network has no structure that distinguishes the node types.
  • Evaluation: The analysis uses all two-event motifs and evaluates the node-type null hypothesis with z-scores.Under the null hypothesis, z-scores should have zero mean and unit variance.
  • Evaluation: The synthetic z-scores do not have zero mean and unit variance, indicating that the null model detects the planted type-dependent pattern.Figure 2 reports this departure from the null expectation.
  • Effect size: r(m) ≈ 1.18 means the overrepresented causal motifs are 18% more common than expected under no node-type effect.The ratio reports effect size, unlike the z-score.

III. RESULTS

The study analyzes a six-month mobile-phone dataset containing 600 million calls among 6.3 million anonymized customers. The records include temporal information for studying motif patterns.

  • Dataset: The mobile-phone dataset contains 600 million calls over six months between 6.3 million anonymized customers.The dataset is used to analyze temporal patterns in communication.

A B C

The analysis uses two-event temporal motifs and examines colored variants based on node types. Synthetic data show that most motifs have near-zero z-scores, while causal chains can produce strong positive or negative deviations.

  • A B C: Most synthetic motifs have z ≈0 because node colors do not generally produce differences.The distribution is averaged over 50 synthetic data sets.
  • A B C: Causal chains with same-color first events produce a peak near z ≈17.The peak corresponds to four causal chains.
  • A B C: Causal chains with different-color first events produce a peak near z ≈−17.This peak contains the four remaining causal chains.
  • A B C: The study analyzes two-event motifs because larger motifs are computationally and analytically more demanding.With 24 node types, there are already 56448 different two-event motifs.

A. Node types affect motif counts

Empirical motif counts depend on node types beyond what the aggregate network predicts. Shuffling node types or event times removes the corresponding differences, supporting the null-model comparison.

  • A. Node types affect motif counts: About 35% of empirical motifs have |z| > 1.96, rejecting independence between motif counts and node types.The observed z-score distribution has neither zero mean nor unit variance.
  • A. Node types affect motif counts: Shuffling node types makes the z-score distribution consistent with the null hypothesis while preserving untyped temporal subgraphs.Randomized node types cannot retain differences between node types.
  • A. Node types affect motif counts: Shuffling event times preserves the aggregate network but removes temporal correlations that could explain differences between node types.The time-shuffled data has exactly the same aggregate network as the original data.

B. There is temporal homophily

The paper tests whether similar people participate jointly in contact sequences more often than aggregate-network structure predicts. Temporal homophily is strongest in more complex motifs and especially for payment-type similarity.

  • B. There is temporal homophily: Temporal homophily means similar individuals jointly participate in interaction patterns beyond homophily observed in the aggregate network.The analysis concerns contact sequences rather than only aggregate ties.
  • B. There is temporal homophily: For two-person motifs, the only significant difference is under-expression of returned calls among participants sharing payment type.Frequent prepaid-to-postpaid calls followed by immediate callbacks explain this pattern.
  • B. There is temporal homophily: Chains and stars are more common when all participants are similar across age group, gender, and payment type.Star motifs are also more common when participants agree on only one attribute.
  • B. There is temporal homophily: Similarity in payment type shows the strongest temporal-homophily effect, although it may correlate with socioeconomic factors.Text-message results are qualitatively similar.

C. Chains and stars are over-expressed for females

Gender differences appear in temporal motif frequencies: all-female motifs are over-expressed for most motifs beyond repeated and returned calls, while all-male motifs are slightly under-expressed.

  • C. Chains and stars are over-expressed for females: All-female motifs are over-expressed for every analyzed motif except repeated and returned calls.The comparison uses motifs whose participants are exclusively female or exclusively male.
  • C. Chains and stars are over-expressed for females: All-male motifs are slightly under-expressed for motifs other than repeated and returned calls.No gender difference is observed for repeated and returned calls.
  • C. Chains and stars are over-expressed for females: The authors describe the gender pattern as consistent with findings that men’s phone use is more instrumental and women’s more conversational.They do not claim a single explanation for the observation.

D. Local edge density correlates with temporal motifs

Temporal motif frequencies vary systematically with local edge density and node attributes, revealing patterns that extend beyond aggregate edge weights. Dense regions contain more multi-edge group interactions, while sparse regions favor repeated and returned calls.

  • Local edge density is used to compare temporal patterns, with dense events inside 4-clique communities and sparse events on all other edges.The analysis distinguishes event types according to whether their aggregate-network edge belongs to a 4-clique community.
  • Repeated and returned calls are more common on sparse edges, whereas other two-event motifs are relatively more common in dense network regions.This ordering remains robust across most combinations of node types.
  • Table I compares mean r(m) for attribute-homogeneous motifs against other motifs across age, gender, payment type, and their combination.A larger first value indicates homophily, with statistical significance assessed using Welch’s t-test and corrected for multiple comparisons.
  • Gender-dependent homophily analysis compares all-female star and chain motifs with corresponding all-male motifs using average r(m).The table also contrasts same-gender motifs with all other motifs and applies the same statistical testing framework.
  • Dense edges are associated not only with higher weights but also with group talk involving more than two individuals.This extends the comparison beyond Granovetter’s weight-based hypothesis to temporal interaction structure.

IV. DISCUSSION

The discussion frames temporal motifs as a way to study group interactions that static networks miss, while highlighting rich temporal structure in mobile-phone data. The framework yields differences independent of aggregate-network structure but remains constrained by its event-overlap assumptions.

  • Temporal motifs assess meso-scale group interactions that cannot be observed in a static network representation.The discussion places these patterns alongside individual-level communication dynamics and system-wide effects such as burstiness.
  • Mobile-phone data show rich meso-scale temporal structure, including robust but not always easily explained correlations such as recipients’ age in out-stars.The connection between temporal motifs and local edge density is also reported as robust.
  • The current framework applies only when nodes have at most one event at a time or when events have no duration.This is identified as the framework’s largest constraint.
  • Temporal motifs provide a framework that can extend beyond social systems whenever time-resolution data are available.The discussion presents this as a general applicability claim rather than a result restricted to mobile communication.
  • The framework’s temporal differences are independent of the aggregate network and therefore complement information from weighted aggregate networks.This makes temporal results distinct from patterns already captured by aggregate connectivity and edge weights.

Appendix A: Mobile phone data

The study analyzes six months of mobile communication data using colored temporal motifs and a null model that controls for aggregate-network structure. It finds attribute-specific temporal patterns and robust differences between dense and sparse network regions.

  • Mobile phone data: The dataset contains 625 million calls and 207 million SMS messages across six consecutive 30-day periods.The analysis is repeated separately by month to assess temporal consistency.
  • Node attributes: Node types combine gender, six age intervals, and prepaid or postpaid payment status, yielding 24 categories for 6.22 million users.The 10-minute time window is intended to allow intentional reactions while limiting coincidental simultaneity.
  • Null model: The null-model analysis preserves aggregate-network structure while sampling motif counts from distributions conditioned on edge-weight sequences.Node colors are incorporated only when constructing the final color-specific motif samples.
  • Attribute patterns: Temporal motif frequencies vary by node attributes, including payment-type effects, gender patterns, and stronger homophily in complex chains and stars.Repeated calls are especially common between prepaid users, while chains and stars show broader similarity patterns across attributes.
  • Local edge density: One-edge motifs are more common on sparse edges, whereas two-edge motifs are more common on dense edges.For two-edge motifs, both-dense configurations are most common, both-sparse configurations follow, and mixed-density configurations are least common.
  • Combined effects: Node-type and edge-density effects are largely independent, as illustrated by returned contacts being more common for a specific payment-and-gender combination and on sparse edges.For prepaid female to postpaid female returned contacts, the reported ratios are 1.048 on sparse edges and 0.898 on dense edges.
Loading 1302.2563v1…