Source-linked AI summary

Using Gossips to Spread Information: Theory and Evidence from a Randomized Controlled Trial

Abhijit Banerjee, Arun G. Chandrasekhar, Esther Duflo, Matthew O. Jackson

arXiv:1406.2293v6physics.soc-phcs.SI

TL;DR

The paper asks whether communities can identify effective information diffusers without network data. It develops a gossip-based model, tests nominations in Indian villages, and finds that seeding nominated individuals increases information spread by more than 65%.

  • Problem

    The paper asks how to identify highly central information diffusers without costly network data, since status labels and friend counts can fail to find diffusion-relevant centrality.

  • Method

    The paper combines a model of gossip-based information flow with village nomination surveys and a randomized field experiment testing nominated individuals as diffusion seeds.

  • Results

    More than a 65% increase in information spread occurred when at least one villager-nominated gossip seed was hit, relative to random or status-based seeds.

  • Takeaways & Limitations

    Simple nominations can provide a cost-effective way to identify effective information-diffusion seeds in reasonably sized communities.

  • Takeaways & Limitations

    The paper’s model excludes trust and endorsement, and its experiments cover communities of roughly one thousand people.

Abstract

from arXiv · show

Is it possible to identify individuals who are highly central in a community without gathering any network information, simply by asking a few people? If we use people's nominees as seeds for a diffusion process, will it be successful? We explore these questions theoretically, via surveys, and via field experiments. We show via a model of information flow how members of a community can, just by tracking gossip about others, identify highly central individuals in their network. Asking villagers in rural Indian villages to name good seeds for diffusion, we find that they accurately nominate those who are central according to a measure tailored for diffusion - not just those with many friends or in powerful positions. Finally, we run a randomized field experiment in 213 other villages that tests how effective it is to use such nominations as seeds for a diffusion process. Relative to random seeds or those with high social status, hitting at least one seed nominated by villagers leads to more than a 65% increase in the spread of information.

1. Introduction

The paper asks whether people can identify diffusion-central individuals without network data and tests whether their nominees improve information spread. Theory, village surveys, and a randomized field experiment show that gossip nominations identify effective seeds, although the mechanism is not fully explained by measured diffusion centrality.

  • Motivation: The paper addresses how to identify highly central information diffusers cheaply, without collecting network data or relying on status labels or friend counts.The motivation is that diffusion requires specific centrality measures, while gathering the relevant network information can be costly and time consuming.
  • Theory: A simple gossip-counting model shows that listeners can estimate others’ diffusion centrality by tracking how often each person is mentioned as an information source.The model gives a possibility result rather than the only possible learning mechanism.
  • Survey evidence: Villagers nominated people who ranked in the top quartile of diffusion centrality, with many nominees ranking in the top decile.Nominations remained correlated with diffusion centrality after controlling for leadership status and geographic position.
  • Field experiment: 3.78 more calls, a 65% increase, resulted when at least one gossip nominee was seeded compared with villages where no gossip nominee was seeded.The experiment covered 213 villages and compared gossip-based seeding with random and village-elder seeding.
  • Field experiment: Gossip nominees produced about twice as many entries as village elders or random non-nominated villagers.The reported comparison concerns information diffusion despite moderate call-back rates.
  • Mechanism: Measured diffusion centrality did not explain all of the extra diffusion from gossip nominees, possibly because nominations capture listening, charisma, talkativeness, or centrality more accurately than the network measure.The coefficient for gossip nomination remained substantial when both nomination and measured diffusion centrality entered the regression, though it became less precise.
  • Contribution: The paper distinguishes its approach from using the friendship paradox to find high-degree individuals: it targets centrality measures more complex than degree centrality.The authors present the nomination process as a cheaper alternative or complement to detailed network data for information diffusion.
  • Scope and limitations: The study focuses on pure information transmission, where participation does not require trust, and experiments communities of roughly one thousand people.The authors caution that trust may matter when diffusion also requires endorsement and that nomination ability may not scale to much larger networks.

2. A Model of Network Communication

The model describes stochastic information transmission over a network and defines diffusion centrality from the sender’s perspective and network gossip from the receiver’s perspective. Its theory links gossip-based reception counts to centrality while emphasizing dependence on transmission rates, time horizons, and network structure.

  • Network setup: The model represents a society of n individuals connected by a possibly directed and weighted adjacency matrix g.The network is assumed to be strongly connected unless otherwise stated.
  • Diffusion process: Information originating at node i is broadcast through neighbors for T periods, with each informed node transmitting independently with probability q.The finite horizon allows information relevance, attention, or conversation to end over time.
  • Diffusion centrality: Diffusion centrality measures the expected total number of times information originating from a node is heard across society during the T-period process.It is based on the hearing matrix, whose ij entry is the expected number of times j hears information originating from i.
  • Relation to other measures: Diffusion centrality spans degree, Katz–Bonacich, and eigenvector centrality as T and q vary.At T = 1 it is proportional to out-degree centrality; under q < 1/λ1 and T = ∞ it coincides with Katz–Bonacich centrality.
  • Parameter regimes: The threshold q = 1/λ1 determines whether the infinite-horizon diffusion-centrality sum converges or diverges, while finite-horizon behavior also changes with T relative to graph diameter.The authors propose q = 1/E[λ1] and T = E[Diam(g)] as a natural empirical benchmark.
  • Network gossip: Network gossip is defined from the receiver’s perspective by tracking information about different sources that reaches a listener through repeated conversations.Listeners can compare cumulative mentions of sources over many topics originating throughout the society.

3. Relating Diffusion Centrality to Network Gossip

The paper shows theoretically that people can infer diffusion centrality from how often they hear gossip originating from others, even without observing network structure. Over time, rankings converge to diffusion centrality under stated network and transmission conditions, while several assumptions limit interpretation.

  • Identifying Central Individuals: Network gossip rankings are positively correlated with diffusion centrality for any q and T.The result uses NG_j, the amount of gossip that j has heard about others.
  • Identifying Central Individuals: Theorem 1 states that the covariance between diffusion centrality and aggregate network gossip equals the variance of diffusion centrality.Thus, when individuals differ in diffusion centrality, average covariance is positive.
  • Identifying Central Individuals: As gossip continues over extended periods, every individual can eventually rank others’ centralities perfectly, including cardinally rather than only ordinally.The result concerns learning through repeated information exchange.
  • Identifying Central Individuals: When q ≥1/λ1 and the network is aperiodic, each listener’s gossip ranking converges to a quantity proportional to diffusion centrality and eigenvector centrality.The convergence is stated as T →∞.
  • Assumptions and Limitations: More sophisticated inference of network topology could accelerate learning, but the result establishes learning using only source tracking and count memory.The model allows weighted and directed networks, with a corresponding eigenvalue condition.
  • Assumptions and Limitations: The learning result requires q ≥1/λ1; as q approaches zero, people hear about others too infrequently for rankings to avoid network-distance effects.The model also assumes information travels through network edges and agents do not assess trust or endorsement.

4. Evidence: who are the gossips?

Survey evidence from rural Indian villages shows that villagers’ nominations identify people who are especially central for diffusion, rather than merely socially prominent or locally connected. Nominations are concentrated, relate to network distance, and remain associated with diffusion centrality after controlling for other characteristics.

  • Data Collection: Survey data combine detailed network information from 33 rural Karnataka villages with villagers’ nominations of good information diffusers.The network data cover 12 types of daily interaction and 89.14% of 16,476 households.
  • Data Collection: Only half of households answered the gossip questions, while conditional on naming someone, nominations were highly concentrated.The passages attribute nonresponse possibly to uncertainty or reluctance to single someone out.
  • Data Collection: Gossip nominees and village leaders largely differ: 86% were neither, 1% were both, 3% were nominees only, and 11% were leaders only.Under the event question, 8% of leaders were nominated, while 25% of nominated gossips were leaders.
  • Who Are the Gossips?: Individuals nominated as both gossips and leaders are more central than gossip-only nominees, who are more central than leader-only individuals.The corresponding centrality distributions exhibit the stated ordering.
  • Who Are the Gossips?: Fewer than 13% of nominations went to someone in the nominator’s direct neighborhood, while over 28% came from a network distance category with more nodes.Closer nominees nevertheless had higher average diffusion-centrality percentiles than nominees at greater distances.
  • Regression Analysis: A one-standard-deviation increase in diffusion centrality is associated with a 0.607 log-point increase in nominations, significant at the 1% level.Degree, eigenvector centrality, and leadership also predict nomination, whereas geographic centrality does not.
  • Regression Analysis: After controlling for other centrality measures and demographic characteristics, diffusion centrality remains the key predictor of nomination.LASSO selects diffusion centrality alone for event nominations, while the evidence distinguishes it less sharply from eigenvector centrality.

5. Experiment: Do gossip nominees spread information widely?

The field experiment tests whether villagers’ gossip nominations identify effective information-diffusion seeds, comparing them with village elders and randomly selected households. Gossip nominees produced substantially more diffusion, although diffusion centrality explains only part of their advantage.

  • 5. Experiment: Do gossip nominees spread information widely?: The primary outcome was the number of calls from unique households, interpreted as participation after hearing about the promotion.The promotion used free missed calls, and prizes were awarded non-rivalrously, reducing incentives to withhold information.
  • 5. Experiment: Do gossip nominees spread information widely?: Gossip villages displayed more large diffusion events than villages seeded randomly or with elders, and the treatment effect significantly increased median calls by 122%.The distribution of calls in gossip villages stochastically dominated the elder and random distributions.
  • 5. Experiment: Do gossip nominees spread information widely?: Hitting at least one gossip seed increased calls by 3.79, a 65% increase relative to villages where no gossip seed was hit.The instrumental-variables estimate was larger but statistically indistinguishable from the OLS estimate and less precise.
  • 5. Experiment: Do gossip nominees spread information widely?: Gossip nomination’s diffusion advantage was only partly accounted for by diffusion centrality, suggesting nominees capture additional characteristics relevant to information spreading.The authors mention factors such as altruism and interest in the information, but standard errors do not identify their precise contribution.

6. Conclusion

The paper concludes that simple gossip observations can identify globally diffusion-central individuals and that nominated seeds outperform elders and other individuals in spreading information. Because nominations are easy to collect, they may provide a cost-effective way to improve diffusion, while their applicability to more consequential information remains an open question.

  • 6. Conclusion: Simple counting in gossip allows even myopic, non-Bayesian agents to identify globally central individuals without knowing the network structure.Villagers did not merely name locally central people; their nominees were globally central within the village.
  • 6. Conclusion: Nominated individuals were more effective at diffusing a simple piece of information than other individuals, including village elders.The conclusion links this result to the field experiment’s direct test of gossip-based seed selection.
  • 6. Conclusion: The model emphasizes network communication mechanics, but villagers appear to incorporate characteristics beyond network position when choosing effective diffusers.Nominees were described as more successful than the average highly central individual.
  • 6. Conclusion: Gossip nominations are easy to collect and can be used alone or with other simple data to identify effective information-diffusion seeds.The authors describe this protocol as a potentially cost-effective way to improve diffusion and outreach.
  • 6. Conclusion: Whether people can identify trusted individuals, especially for advice such as immunization, remains an open research question beyond the cellphone-giveaway setting.The paper contrasts innocuous giveaway information with potentially more consequential advice.

Tables

The tables document the datasets, nomination predictors, and experimental outcomes used to study gossip-based seed selection and information diffusion.

  • Factors predicting nominations: Tables 3 and 4 report factors predicting the expected number of nominations using network centrality and demographic characteristics.The specifications include degree, eigenvector centrality, diffusion centrality, leadership, geographic position, caste, and village fixed effects; Table 4 also uses post-LASSO selection.
  • Calls received by treatment: Table 5 reports calls received by treatment, including reduced-form, OLS, and instrumental-variables estimates.The outcome is the number of calls received, also normalized by the randomly assigned number of seeds.
  • Calls received by seed type: Table 6 reports calls received according to the characteristics of the selected seeds.The analyses distinguish seeds above the mean by one standard deviation in diffusion centrality and control for the number of gossips, elders, and seeds.

Appendix A. Threshold Parameters (q, T) for Diffusion Centrality

Appendix A characterizes diffusion centrality in large random networks and identifies threshold values of transmission probability and diffusion duration separating limited from expanding diffusion.

  • Theoretical setup: The appendix develops theoretical results showing that diffusion centrality has distinct intermediate parameter values rather than only reducing to boundary centrality measures.Theorem A.1 and Corollary A.1 formalize these properties.
  • Expected diffusion centrality: Theorem A.1 gives the expected diffusion centrality of a typical node in an Erdős–Rényi network when T = o(pn).The expression describes average diffusion in large graphs, while realized nodes can differ in centrality.
  • Threshold parameters: 1/E[λ1] is the threshold for q separating diffusion that reaches a vanishing number of nodes from diffusion that reaches an expanding number.If q = o(1/E[λ1]), expected diffusion centrality tends to 0; if 1/E[λ1] = o(q), it tends to infinity.
  • Threshold parameters: E[Diam(g(n, p))] is the threshold for T separating diffusion that reaches a vanishing fraction from diffusion that reaches a full fraction of nodes.The critical values q = 1/E[λ1] and T = E[Diam(g)] mark transitions between diffusion regimes.
  • Threshold parameters: At the critical value itself, diffusion reaches a non-trivial fraction of the network but not everyone.The appendix describes this as the boundary between diffusion reaching almost nobody and diffusion saturating the network.
  • Parameter choice: Choosing q = 1/E[λ1] and T = E[Diam(g(n, p))] removes free parameters from diffusion centrality for comparisons with other measures.This parameterization is presented as an interesting centrality measure distinct from standard measures at these values.

B.1. Relation of Diffusion Centrality to Other Measures.

This section relates diffusion centrality to degree, Katz–Bonacich, and eigenvector centrality by varying the communication parameter q and time horizon T.

  • Network setting: The theoretical comparison is developed for weighted and directed networks, allowing heterogeneous and nonreciprocal communication links.The framework retains q explicitly even though it is redundant for a fully heterogeneous communication matrix.
  • Boundary cases: With one period, diffusion centrality is proportional to out-degree centrality.This is the short-horizon boundary case.
  • Boundary cases: When q ≥ 1/λ1 and T approaches infinity, diffusion centrality approaches eigenvector centrality.The limiting ranking is determined by the first right-hand eigenvector, up to ties.
  • Boundary cases: With T = ∞ and q < 1/λ1, diffusion centrality coincides with Katz–Bonacich centrality.Limited diffusion persists in this parameter regime.

B.2. Other Proofs.

The proofs establish the large-network formulas, threshold results, and limiting relationships that support the appendix’s characterization of diffusion centrality.

  • Proof strategy: The proof of Theorem A.1 bounds repeated walks in Erdős–Rényi networks to derive the asymptotic expected diffusion centrality.The argument uses the condition T << pn and properties of expected network powers.
  • Proof of threshold results: The threshold proof substitutes npq into the theorem’s expression to obtain the vanishing-versus-expanding diffusion result.Monotonicity in q extends the lower-q conclusion from the threshold argument.
  • Proof of threshold results: The proof of the duration threshold compares diffusion growth with network size and uses the relationship between network diameter and logarithmic distance scales.The argument distinguishes durations below and above the relevant logarithmic threshold.
  • Proof of eigenvector limit: The limiting proof shows that higher eigenvalue terms vanish relative to the leading eigenvalue when q exceeds the spectral threshold.Consequently, columns of the limiting matrix become proportional to the first right-hand eigenvector, preserving its ranking up to ties.

Appendix C. Extension of Phase 1 results

Appendix C extends the Phase 1 nomination analysis by replacing Poisson specifications with OLS and adding post-LASSO estimation.

  • The appendix repeats the Phase 1 descriptive analyses using OLS specifications instead of Poisson specifications.
  • Post-LASSO selects variables explaining the number of nominations before recovering consistent parameter estimates.
  • The appendix reports OLS regressions for expected nominations under both the event and loan questions.
  • The specifications include degree, eigenvector, and diffusion centrality normalized by their standard deviations.

Appendix D. Extension of experiment analysis

Appendix D extends the experiment analysis by using four instruments and reporting calls received as the primary outcome, including treatment and gossip-hit specifications.

  • The appendix extends the experiment results by using four instruments.
  • The analysis uses the number of calls received as its primary outcome.
  • The appendix also normalizes calls received by the randomly assigned number of seeds, 3 or 5.
  • Reduced-form specifications regress calls on gossip-treatment and elder-treatment indicators, while another specification uses whether at least one gossip was hit.

Appendix E. Experiment Analysis with Broadcast Village

Appendix E repeats the main experimental analyses while including the broadcast village where a seed made the poster, and examines calls by treatment and seed characteristics.

  • The appendix repeats the main experimental analyses after including the broadcast village.
  • The broadcast village is included when the poster was made by one of the seeds.
  • The appendix includes tables reporting calls received by treatment and by seed type.
  • The analysis reports calls received and calls normalized by the randomly assigned number of seeds, 3 or 5.
  • OLS regressions relate calls received to seed characteristics while controlling for total gossips, elders, and seeds.
Loading 1406.2293v6…