Source-linked AI summary

Respondent-Driven Sampling: An Assessment of Current Methodology

Krista J. Gile, Mark S. Handcock

arXiv:0904.1855v1stat.AP

TL;DR

RDS provides a way to sample hard-to-reach populations, but its current estimators depend on assumptions whose properties are not fully understood. This paper uses simulations that retain branching and without-replacement features to evaluate sensitivity to seeds, respondent behavior, and sampling assumptions, finding that estimator performance is condition-dependent and warrants caution.

  • Problem

    Existing RDS estimators rely on approximations and assumptions whose properties have not been systematically explored, despite the method’s use for hard-to-reach populations.

  • Method

    The paper uses computational simulations of RDS sampling while retaining its branching and without-replacement features.

  • Results

    Current RDS estimators are sensitive to seed selection, respondent behavior, and deviations from the random-walk sampling model.

  • Takeaways & Limitations

    RDS can reach large and varied samples of hard-to-reach populations, but its estimators’ favorable statistical properties depend heavily on often-unrealistic assumptions.

  • Takeaways & Limitations

    The current estimator uses draw-wise probabilities from a with-replacement process to approximate list-wise inclusion probabilities in typically without-replacement sampling.

Abstract

from arXiv · show

Respondent-Driven Sampling (RDS) employs a variant of a link-tracing network sampling strategy to collect data from hard-to-reach populations. By tracing the links in the underlying social network, the process exploits the social structure to expand the sample and reduce its dependence on the initial (convenience) sample. The primary goal of RDS is typically to estimate population averages in the hard-to-reach population. The current estimates make strong assumptions in order to treat the data as a probability sample. In particular, we evaluate three critical sensitivities of the estimators: to bias induced by the initial sample, to uncontrollable features of respondent behavior, and to the without-replacement structure of sampling. This paper sounds a cautionary note for the users of RDS. While current RDS methodology is powerful and clever, the favorable statistical properties claimed for the current estimates are shown to be heavily dependent on often unrealistic assumptions.

1 Introduction to Respondent-Driven Sampling

RDS uses link-tracing recruitment to reach hard-to-reach populations and reduce dependence on convenience-selected seeds. This paper motivates a systematic simulation-based evaluation because existing estimators rely on incompletely tested assumptions.

  • 1 Introduction to Respondent-Driven Sampling: RDS addresses populations that are rare, stigmatized, or absent from usable sampling frames, making standard probability sampling difficult or prohibitively expensive.Alternative approaches may be costly, conditional on times and locations, or unable to represent marginalized population members adequately.
  • 1 Introduction to Respondent-Driven Sampling: Current RDS inference reduces dependence on the initial convenience sample through many waves and estimates inclusion probabilities using a Markov Chain representation.This estimation strategy was proposed by Salganik and Heckathorn and extended by Volz and Heckathorn.
  • 1 Introduction to Respondent-Driven Sampling: RDS follows social-network links, using respondents to recruit alters through successive sampling waves.The design typically gives respondents a fixed number of coupons to distribute among other members of the target population.
  • 1 Introduction to Respondent-Driven Sampling: The paper evaluates estimator sensitivity to seed selection, respondent behavior, and the with-replacement sampling assumption.It emphasizes without-replacement sampling as a relatively underexamined issue because the process is branching and difficult to capture analytically.
  • 1 Introduction to Respondent-Driven Sampling: Because RDS combines branching, without-replacement recruitment, arbitrary social graphs, and convenience-selected seeds, the study relies on computational simulations to retain these features.The authors report that analytical approximations were either intractable or unable to model critical aspects of the design.

2 Overview of Existing RDS Estimation

RDS estimation uses link-tracing data and Markov-chain-inspired estimators to reduce dependence on convenience-sample seeds. The paper argues that these estimators rely on strong assumptions about mixing, respondent behavior, replacement, and uncertainty estimation.

  • RDS estimation framework: RDS estimation leverages successive sampling waves to reduce dependence on the initial convenience sample.This strategy is intended to increase the validity of inference for hidden populations.
  • RDS estimation framework: The V-H estimator uses estimated inclusion probabilities in a generalized Horvitz-Thompson form and generally out-performs the S-H estimator in the simulations.Most results in the paper therefore concern the V-H estimator, aside from the direct comparison in Section 4.
  • The current RDS estimator: The V-H approximation treats sampling as a random walk whose stationary draw-wise probabilities are proportional to node degree.This model assumes an undirected connected network and random referral among alters.
  • The current RDS estimator: Early transition probabilities depend on seed values, while later probabilities can converge toward starting-node independence.The paper visualizes this convergence with colored transition matrices across random-walk steps.
  • Assumptions of the current estimator: The model omits important features of actual RDS, including branching, without-replacement sampling, non-equilibrium starts, and respondent-behavior assumptions.These deviations create multiple sources of potential estimator bias.
  • Assumptions of the current estimator: Current inference requires a small sample fraction for its with-replacement approximation, while sufficiently many waves are needed for convergence and seed-bias reduction.These requirements can conflict, especially when network structure slows mixing.
  • Estimators of uncertainty: The bootstrap uncertainty procedure can retain seed-driven overrepresentation under strong homophily, and the alternative variance estimator has only been studied in a simpler with-replacement setting.The paper therefore questions whether current uncertainty estimates provide reliable coverage in realistic RDS settings.

3 Simulation Study to Evaluate Sensitivity to Assumptions

The simulation study isolates how RDS design and implementation features affect estimator performance, using realistic parameters drawn largely from CDC pilot data.

  • 3 Simulation Study to Evaluate Sensitivity to Assumptions: The simulations isolate effects of selected RDS study-design and implementation features.Parameters were chosen to match CDC surveillance pilot-data characteristics as closely as possible.

1. Simulate 1000 networks with a specified structure

Each network is subjected to an RDS sampling-process variant with specified characteristics, enabling controlled simulation of the sampling design.

  • 1. Simulate 1000 networks with a specified structure: The study implements an RDS sampling-process variant with desired characteristics on each network.

4. Compare each estimate to the known true population proportion

The simulations generate controlled networked populations and RDS samples, then evaluate estimator sensitivity by comparing estimates with known population proportions under varied structural and behavioral conditions.

  • 4. Compare each estimate to the known true population proportion: The study repeats simulations across variants of network structure and the sampling process.
  • 4. Compare each estimate to the known true population proportion: The experiments examine wave sufficiency, homophily, later-wave estimation, non-random referral, and deviations from with-replacement sampling.
  • 4. Compare each estimate to the known true population proportion: The baseline simulations use populations of 1000, samples of 500, and mean degree 7.These settings approximate CDC surveillance targets and pilot-data characteristics.
  • 4. Compare each estimate to the known true population proportion: Infection status is assigned to 20% of simulated population members as the discoverable characteristic for estimating hidden-population proportions.
  • 4. Compare each estimate to the known true population proportion: ERGMs generate networks with prespecified population size, homophily, relative activity levels, and mean degree.The network structure is determined by selected network statistics and their parameters.
  • 4. Compare each estimate to the known true population proportion: Most simulations set infected-infected edge probability at five times the mixed-dyad probability and uninfected-uninfected probability at twice the mixed-dyad probability.These settings maintain a constant mean degree across groups with 20% infected members.
  • 4. Compare each estimate to the known true population proportion: Relative activity levels of infected and uninfected nodes are varied in some substudies while mixed and infected-infected edge probabilities remain proportionally fixed.Unless otherwise noted, both groups have equal activity levels.
  • 4. Compare each estimate to the known true population proportion: The ERGM represents mean degree, relative activity levels, and infection-status homophily through network statistics implemented in statnet.

3.1 Removing Seed Bias

Removing seed bias depends on sufficient mixing, which is hindered by homophily and can remain substantial with typical RDS wave counts. Discarding early waves may reduce bias in some settings but can also reverse or amplify it because sampling is without replacement.

  • Sufficiently Many Waves of Sampling: Typical RDS wave counts may be insufficient for the nodal mixing required to obtain asymptotically unbiased estimators.The effect depends strongly on population clustering and the number of waves.
  • Sufficiently Many Waves of Sampling: With random seed selection, six-wave and four-wave samples show little appreciable bias and similar variance because seeds approximate the theoretical equilibrium distribution.The comparison uses 6 seeds across 6 waves versus 20 seeds across 4 waves, with samples of 500.
  • Sufficiently Many Waves of Sampling: Biased seeds produce substantially lower bias with fewer waves than with more waves, because shorter chains increase dependence on the seeds.This pattern occurs for both infected and uninfected seed simulations.
  • Homophily Weak Enough: 3.5 times the variance occurs under higher homophily, which also produces 1.9 times the standard deviation under unbiased seeds.Higher correlation among successive sampled infection statuses provides less information at comparable sample sizes.
  • Homophily Weak Enough: Higher homophily produces far greater bias under biased seed selection by increasing dependence between seed characteristics and subsequent samples.Transition probabilities remain clustered even after 14 steps, so seed-induced bias persists.
  • Estimation based on later waves only: Discarding the seeds and first wave makes bias virtually disappear with little variance impact, but discarding two or three waves can restore or amplify bias in opposite directions.The reversal occurs because discarded nodes cannot be re-sampled in the without-replacement process.

3.2 Respondent Behavior: Random Referral

The study examines whether respondent referral behavior affects V-H estimates. Referral that favors infected alters increases positive bias across all seed-selection conditions because the estimator does not account for this preference.

  • Random Referral: Referral making infected alters 20% more likely to be sampled increases positive bias under uninfected, random, and infected seed selection.The increased bias has about the same magnitude across the three seed-selection regimes.

3.3 Random Walk Model

The V-H estimator is evaluated against violations of its random-walk assumptions, especially large sample fractions and with-replacement modeling. Bias can become substantial when activity levels differ, while without-replacement sampling can improve performance relative to the theoretical approximation.

  • Large sample fractions: At equal activity (w = 1), estimator bias is negligible across population sizes, while variance decreases as the sample fraction increases.Sampling 95% of the population leaves little room for estimator variance.
  • Mechanism of bias: The bias persists because near-uniform true inclusion probabilities conflict with degree-proportional estimator weights, under-representing infected nodes when they have higher degree.The full-population limit has uniform inclusion probabilities, whereas the estimator continues down-weighting higher-degree nodes.
  • Activity differences: As infected-node activity rises to w = 1.5, 1.8, and 3, negative bias increases and becomes more severe at larger sample fractions.With w = 3, even the largest estimates are more than .03 below the true value at 50%; at 95%, estimates remain below half the true proportion.
  • Implications: The sampling-weight error may be correctable, with corrections proposed by Gile (2008) and Gile and Handcock (2009).The simulations also found no substantial differences when varying infection-status homophily or clustering on an orthogonal variable.
  • With-replacement comparison: With-replacement simulations produce greater bias and variance than without-replacement simulations, particularly when seeds are biased.Without-replacement sampling accelerates mixing, whereas with replacement can repeatedly sample the same subgroup.

4 Comparison of the Classic and Current RDS Estimators

The paper compares the classic Salganik-Heckathorn estimator with the current Volz-Heckathorn estimator. Across nearly all simulated circumstances, the V-H estimator performs better, although the comparison inherits substantial assumptions from the classic estimator.

  • Classic estimator: The S-H estimator relies on estimated group degrees and cross-group referral proportions derived from generalized Horvitz-Thompson estimates and observed referrals.Its construction equates estimated relations from group A to B with those from B to A.
  • Classic estimator: The classic estimator depends on many assumptions, some acknowledged as approximations, and substitutes estimates into its equations in an ad hoc manner.The paper therefore presents only a brief comparison of the two estimators.
  • Overall comparison: Nearly all relative efficiencies of S-H to V-H are below 1, indicating superior V-H performance.In nearly all such cases, V-H has both lower bias and lower variance.
  • Exception: In the biased-referral, all-infected-seeds case, S-H has relative efficiency 2.72 because referral bias partly offsets seed-selection bias.This is identified as a notable exception to V-H’s generally superior performance.
  • Overall comparison: The V-H estimator outperforms the S-H estimator in almost all circumstances and is easier to compute.It also applies directly to continuous as well as categorical variables.

5 Discussion and Recommendations

RDS can reach varied hard-to-reach samples, but its inferential performance is sensitive to seed selection, network structure, respondent behavior, and large sampling fractions. The authors recommend foundational research and cautious study design before relying on current estimators.

  • Overall assessment: Current RDS estimators are sensitive to convenience-selected seeds, respondent behavior, and deviations from the random-walk model.These three areas organize the paper’s assessment of current methodology.
  • Seed selection: Typical numbers of waves may not remove seed bias, especially in highly clustered populations.In finite populations, discarding early waves can introduce additional bias because sampled nodes cannot be re-sampled.
  • Seed selection: Increasing waves from 4 to 6 substantially reduced bias under the simulations, but more waves compete with the goal of reaching all population subgroups.Diverse seeds can reduce estimator variance and improve coverage of population components.
  • Network structure: High homophily increases uncertainty and seed-bias effects, and disconnected networks can make representation depend on seed selection.The authors identify circumstances in which RDS may not be appropriate, or may not be appropriate for the full population.
  • Respondent behavior: Referral behavior, alter reporting, and coupon-passing relationships may depart from the assumptions used by current estimators.The paper recommends surveys, follow-ups, qualitative interviews, or participant observation to investigate these features.
  • Respondent behavior: Reliable degree questions that reflect the coupon-passing network are needed for reliable RDS estimates, while non-reciprocal referrals may require more complex evaluation methods.Identified referral biases could potentially be measured and incorporated into analysis.
  • Sampling fraction: With differential activity and homophily, larger sample fractions can produce significant estimator bias.The bias increases with greater activity differentials, even as sampling covers more of the population.
  • Overall assessment: The paper concludes that current RDS estimators’ favorable statistical properties depend heavily on often-unrealistic assumptions.The sampling strategy remains useful for reaching hard-to-reach populations, but inference requires caution.

APPENDIX

The appendix repeats truncated plots so their full data ranges remain visible. The supplied appendix passages provide plot-range and axis-label information but no substantive comparison or outcome.

  • Appendix: The appendix repeats earlier plots whose values fell outside consistent plotting ranges.The repeated plots display the full ranges of data values.
  • Appendix: The supplied appendix labels include estimated proportion infected and values from 0.10 to 0.30.These labels describe displayed quantities rather than reporting a new result.
  • Appendix: One supplied label identifies homophily conditions as low or high and seed categories as uninfected, random, or infected.The passage provides labels but not an outcome comparison among them.
Loading 0904.1855v1…