Source-linked AI summary
Diagnostics for Respondent-driven Sampling
Krista J. Gile, Lisa G. Johnston, Matthew J. Salganik
TL;DR
RDS enables sampling of hard-to-reach populations, but inference relies on strong assumptions about a partly uncontrolled sampling process. This paper develops practical diagnostics for most of those assumptions and applies them to 12 studies, while identifying interpretive limits and opportunities for further refinement.
Problem
RDS estimates require many assumptions, while the dependence between recruiters and recruits makes standard diagnostic tests difficult to apply.
Method
The paper develops intuitive, often graphical diagnostics that use response timing, repeat visits, and multiple seeds, then applies them across 12 RDS studies.
Results
The diagnostics were applied to 12 RDS studies covering female sex workers, drug users, and men who have sex with men.
Takeaways & Limitations
The diagnostics are intended to help RDS researchers understand their sampling processes and motivate future methodological developments.
Takeaways & Limitations
Diagnostic flags may disagree with experienced researchers’ judgments, and a lack of evidence of a problem does not establish that no problem exists.
Abstract
from arXiv · showhide
Respondent-driven sampling (RDS) is a widely used method for sampling from hard-to-reach human populations, especially groups most at-risk for HIV/AIDS. Data are collected through a peer-referral process in which current sample members harness existing social networks to recruit additional sample members. RDS has proven to be a practical method of data collection in many difficult settings and has been adopted by leading public health organizations around the world. Unfortunately, inference from RDS data requires many strong assumptions because the sampling design is not fully known and is partially beyond the control of the researcher. In this paper, we introduce diagnostic tools for most of the assumptions underlying RDS inference. We also apply these diagnostics in a case study of 12 populations at increased risk for HIV/AIDS. We developed these diagnostics to enable RDS researchers to better understand their data and to encourage future statistical research on RDS.
1 Introduction
RDS provides a practical way to collect information from hard-to-reach populations, but inference depends on strong assumptions and a partly uncontrolled sampling design. This paper develops diagnostics for those assumptions and applies them across 12 Dominican Republic studies.
- 1 Introduction: RDS has been used in more than 120 HIV-related studies across 20 countries and adopted by organizations including the CDC.Its use reflects the need for information about populations such as female sex workers, drug users, and men who have sex with men.
- 1 Introduction: Inference from RDS data requires many assumptions, and recent research has challenged the quality of estimates derived from these data.The assumptions concern the sampling process, the underlying population, and respondent behavior.
- 1 Introduction: RDS uses peer referral through respondents’ social networks to recruit successive waves from hard-to-reach populations.Researchers typically select 5 to 10 seeds, provide coupons, and use incentives to support recruitment.
- 1 Introduction: The paper’s primary focus is developing methods to detect violations of RDS assumptions in practice, complementing analytical assessment and alternative-estimator development.The authors emphasize intuitive, graphical diagnostics that can sometimes be used during data collection.
- 1 Introduction: The Volz-Heckathorn estimator relies on assumptions including a random-walk approximation, reduced seed dependence, reciprocated ties, accurately measured degree, and random referral.The true RDS process is without replacement, although the random-walk model requires with-replacement sampling.
3 Case study: 12 sites in the Dominican Republic
The case study applies RDS diagnostics to 12 studies in the Dominican Republic, covering FSW, DU, and MSM across four cities. The results show frequent finite-population effects, while their impact on estimates varies across populations and indicators.
- Study design: The case study used 12 parallel RDS studies of FSW, DU, and MSM in four Dominican Republic cities.The studies were conducted in spring 2008 using standard RDS methods.
- Study design: 3,866 people participated, and 1,677 (43%) completed a follow-up survey.Follow-up data supported diagnostics involving recruitment attempts and reported prior participation.
- Findings: Three studies failed to attain their intended sample sizes, providing strong evidence that earlier samples affected later sampling decisions.The affected studies were FSW-BA, MSM-BA, and MSM-HI.
- Diagnostic approach: The analyses examined finite-population effects using failed recruitment attempts, reported contacts who had already participated, time trends, and estimator sensitivity.Population sizes were estimated through meta-analysis and a degree-sequence method, then used to compare Successive Sampling and Volz-Heckathorn estimates.
- Findings: Nearly all sites showed finite-population effects on at least one indicator, although different indicators often produced different results.The authors attribute these differences either to random variation or to indicators capturing different features of the underlying process.
- Findings: Comparison of the Volz-Heckathorn and Successive Sampling estimators was the most effective diagnostic of global effects on estimates.Its limitation is that the Successive Sampling estimator requires difficult-to-construct study-population size estimates.
5 Detecting convergence
The paper proposes assessing convergence by tracking how RDS estimates change as observations accumulate, rather than relying on chain-length calculations. Visual and flagging diagnostics can identify lingering seed influence, but stable estimates do not guarantee that the sample has explored the population.
- Seed selection can influence finite-sample RDS estimates despite asymptotic results showing seed effects vanish as sample size approaches infinity.Because RDS seeds are selected as an ad-hoc convenience sample, practical samples may retain seed bias.
- Convergence plots track estimated trait prevalence after each non-seed observation and compare the evolving estimates with the complete-sample estimate.The approach focuses directly on the dynamics of the RDS estimate rather than simulated Markov-chain sampling behavior.
- For one Barahona drug-use trait, the estimate rose sharply from 8% among the first 50 respondents to 67% among the final 50.The same sample produced a stable estimate for unprotected sex during the second half, showing convergence is trait-specific rather than purely sample-specific.
- With τ = 50 and ϵ = 0.02, traits are flagged when any of the final 50 estimates differs from the final estimate by more than 0.02.The procedure was applied to 120 group × trait × city combinations, with the highest flagging rate in MSM data at 37.5%.
- The authors recommend inspecting convergence plots during data collection and collecting more data when estimates remain unstable.If further collection is impossible, researchers may consider estimators designed to address features such as seed bias.
- Convergence plots can miss problems when estimates appear stable before recruitment reaches a previously unexplored population segment.They provide some, but not perfect, confidence about convergence.
6 Detecting bottlenecks
The paper detects RDS bottlenecks by comparing recruitment trees originating from different seeds, treating them as natural experiments for identifying distinct communities. Visual plots and a weighted permutation statistic reveal possible bottlenecks, but non-detection does not establish that none exist.
- Bottlenecks can bias RDS estimates when socially separated communities differ in trait prevalence, even if connections within communities are strong.A street-versus-brothel sex-worker example illustrates how limited cross-community recruitment can make estimates problematic.
- Bottleneck plots display estimate dynamics separately for each seed’s recruitment tree to assess whether trees become trapped in distinct communities.In Santo Domingo MSM data, one of three dominant trees produced a markedly different drug-use estimate, unlike the similar tree estimates for Barahona employment.
- The weighted squared deviation uses tree-specific estimates and tree sample sizes to quantify disagreement among seed-originating recruitment trees.The permutation procedure fixes chain lengths and weights, permutes traits, repeats 10,000 times, and flags unusually large observed deviations.
- 41% of FSW, 30% of MSM, and 23% of DU group × trait × city combinations were flagged for possible bottlenecks.No trait was flagged in all four cities, while likely bottleneck sources varied across populations and traits.
- The authors recommend bottleneck plots during data collection and permutation testing when many plots make manual inspection impractical.Evidence of bottlenecks should prompt consideration of additional data collection because estimates may be unstable.
- Flagging can disagree with expert visual interpretation, and statistical properties remain unknown because the dependence structure is unknown.A bottleneck can also remain invisible when all seeds originate in one community and recruitment never reaches another.
7 Reciprocation
The paper evaluates whether recruiters and recruits would recognize and recruit each other, addressing the reciprocity assumption used in RDS inference. Reciprocation was common overall but varied across populations and sites, motivating direct measurement of both relationship and recruitment willingness.
- The follow-up questionnaire asked whether each coupon recipient would have given the recruiter a coupon without the recruiter participating first.This question directly assesses willingness to recruit in the reverse direction.
- About 88% of responses indicated reciprocation, with higher rates in Santiago and lower rates among DU participants.DU reciprocation was especially low in Higuey and Barahona, where coupon selling was considered possible.
- RDS reciprocity requires that recruiters and recruits know each other and would both be willing to recruit one another.The authors recommend collecting relationship information and directly assessing hypothetical reciprocal recruitment.
8 Measurement of Degree
The paper examines whether self-reported degree measurements are timely, reliable, and consequential for RDS estimates. The one-week frame appeared reasonable, reliability was relatively low, and estimate differences were usually small in absolute terms but sometimes large relative to baseline prevalence.
- Because RDS estimators can depend critically on self-reported degree, the paper develops empirical checks of degree measurement and its effects on estimates.The study assesses the validity of the one-week frame, test-retest reliability, and robustness to inconsistent reporting.
- Degree was measured through four questions narrowing from known drug users to those seen during the past week, with the response to question G used for estimation.Additional questions assessed how many contacts respondents could reach by the next day or following week.
- 92% of respondents’ alters were reported as reachable within one week, supporting the one-week time frame in these studies.The authors recommend checking this time window in future studies involving different populations.
- The median difference between initial and follow-up degree responses was 0, but Spearman rank correlations ranged from 0.17 to 0.47, indicating relatively low reliability.Median correlations were 0.33 for FSW and 0.41 for DU and MSM.
- Differences between disease-prevalence estimates based on initial versus follow-up degree ranged from 0 to 0.08, with a median difference of 0.01.These comparisons used participants who completed both interviews.
- In about one-quarter of cases, the estimate difference exceeded 50% of the original estimate despite generally small absolute differences.The authors note that such relative changes could matter in public-health disease surveillance.
9 Participation Bias
Participation bias can arise at multiple stages of RDS recruitment, acceptance, and study participation, potentially making recruits differ systematically from respondents’ contacts. The paper presents graphical and survey-based diagnostics to detect and monitor these processes, while emphasizing that they do not quantify the resulting impact on estimates.
- Recruitment decisions: Recruitment decisions may depart from simple random selection because recruiters choose whom to coupon, recruits choose whether to accept, and accepted recruits choose whether to participate.These three decisions define the paper’s framework for examining participation bias.
- Recruitment effectiveness: Differential recruitment effectiveness can bias estimates when respondents with a trait recruit at different rates and are connected to others sharing that trait.Among FSW in Higuey, respondents with HIV recruited at only half the rate of those without HIV.
- Recruitment bias: Recruitment Bias Plots compare trait composition among respondents’ contacts, coupon recipients, and recruits to identify systematic changes across referral stages.The analysis restricts attention to recruiters with data at all three levels.
- Recruitment bias: Across every site, reported employment increased from contacts to coupon recipients to recruits, suggesting selective coupon distribution and return among employed contacts.The authors caution that survey response bias, including social desirability, could also explain this pattern.
- Non-response: Non-response occurs when intended respondents refuse coupons or fail to return them, and its rates varied substantially across sites.Coupon refusal exceeded 50% among FSW in Santo Domingo, while total non-response ranged from 62.3% to 26.3% across cited sites.
- Participation motivation: Participation motivations may relate to study outcomes, and such relationships can introduce bias when motivation predicts both participation and the outcome.The odds of having HIV among those motivated by HIV testing ranged from 0.43 to 2.03 times the odds among others across cited MSM sites.
- Implications: These diagnostics monitor potential participation bias and can inform sampling adjustments, estimator choice, or new inferential methods, but they do not directly measure estimate distortion.The authors recommend pairing quantitative diagnostics with qualitative evaluation of recruitment and participation decisions.
10 Discussion
RDS enables sampling from populations that are otherwise difficult to reach, but it trades this practical benefit for high variance and reliance on many assumptions. The paper recommends a staged diagnostic program, combined with qualitative work, to improve understanding of RDS sampling processes and support future methodological development.
- Discussion: RDS converts an initial convenience sample into a dependent link-tracing sample while treating the final sample as a probability sample with known or estimable probabilities.This design differs from traditional surveys whose sampling procedures are conducted within researcher-controlled sampling frames.
- Discussion: RDS is often used because alternative workable strategies are unavailable, despite its large estimate variance and numerous inferential assumptions.The paper identifies these as two main costs of RDS.
- Recommendations: Researchers should include the paper’s analyzed questions in initial and follow-up questionnaires and combine quantitative diagnostics with qualitative analysis when possible.The recommendations are intended for planning future data collection.
- Recommendations: During data collection, recommended checks include convergence, bottleneck, all-points, recruitment-effectiveness, recruitment-bias, reciprocation, and non-response analyses.These diagnostics target traits of interest and recruitment processes.
- Recommendations: After data collection, researchers should examine motivation-outcome relationships, finite-population effects, degree-question validity and reliability, and unweighted sample means.These checks complement diagnostics conducted during sampling.
- Future research: The authors expect continued refinement of RDS diagnostics as knowledge about sampling and estimators develops.They aim to improve researchers’ understanding of sampling processes and spur methodological developments.
S1 With-replacement Sampling
The paper evaluates indicators of finite-population effects in 12 RDS studies using degree trends, population-size estimates, and comparisons of Successive Sampling and Volz-Heckathorn prevalence estimates. Evidence was strongest for MSM-SA and possible for MSM-SD, while most sites showed positive or null degree trends.
- Finite-population indicators: Figures S1 and S2 visualize, by seed and site, how the proportion of respondents’ contacts who had already participated changes over sample time.Within some seeds, periods of low proportions were followed by higher proportions.
- Finite-population indicators: Higher-degree people are expected to be sampled earlier as a population is depleted, so decreasing degree over time can indicate finite-population effects.The analysis compared linear and log-degree trends over study order.
- Finite-population indicators: Robust methods flagged only 1–3 of 12 sites for decreasing degree, compared with 5 of 12 using the non-robust linear model.The authors therefore found little overall evidence of decreasing degree over time.
- Finite-population indicators: MSM-SA showed negative degree trends under both approaches, whereas FSW-BA showed positive trends and MSM-SD produced conflicting trends driven by a few high early responses.Figure S3 illustrates these fitted relationships.
- Finite-population indicators: Robust indicators clearly suggested finite-population effects for MSM-SA and possibly MSM-SD, while the other populations showed positive or null trends.This included populations that had not reached their target sample sizes.
- Estimator comparison: When population-size estimates are available, researchers can compare Successive Sampling and Volz-Heckathorn estimators; here, population sizes were estimated from the RDS data because external estimates were unavailable.The approach required specifying a prior distribution for population size.
- Estimator comparison: The population-size analysis used the posterior mean, the lower posterior highest-probability-density bound, and a 1% lower-bound assumption for MSM populations.Prevalence estimates were compared for traits whose maximum absolute estimator difference exceeded .01.
S2 All Points Plot
The All Points Plot supplements convergence and bottleneck plots by displaying every respondent’s trait value by seed and sample order. In the Higuey MSM example, it helps reveal when a bridge-group subgroup entered the sample and whether its estimate may depend on seed pathways.
- S2 All Points Plot: All Points Plots show respondents’ trait values jointly by seed and sample order, preserving information obscured when convergence and bottleneck plots are viewed separately.Convergence plots obscure differences across trees, while bottleneck plots obscure variation over time.
- S2 All Points Plot: For MSM in Higuey, self-identified heterosexual respondents formed a bridge group relevant to infection spread between the high-risk MSM group and the larger heterosexual population.The example uses this subgroup to demonstrate how the plots work together.
- S2 All Points Plot: The subgroup was absent from the first 100 observations and appeared later, while the estimate p̂ = 0.12 had not clearly stabilized.The convergence pattern raised concern that the final estimate might be influenced by seed choice.
- S2 All Points Plot: The bottleneck plot indicated that self-identified heterosexual respondents were reached only through certain recruitment trees, suggesting possible bottlenecks.The all-points display further showed that these respondents arrived late in the sample.
S3 Reciprocation
The paper evaluates reciprocation across all reported network ties because RDS sampling probabilities depend on network connections, not only coupon-passing ties. Among 3,860 respondents, reported ability to give and receive coupons often differed, motivating normalized comparisons.
- Reciprocation: RDS estimators require reciprocated network ties because sampling probabilities relate to incoming connections, while respondents more readily report outgoing connections.If ties are reciprocated, self-reported out-degree equals in-degree, making reported network size informative for sampling probabilities.
- Measurement: The study measured whether respondents could receive coupons from contacts and give coupons to contacts within specified time frames.Questions Q–S asked about contacts who could give a coupon within a week, contacts respondents could give coupons to, and contacts reachable by the following week.
- Reciprocation: Among 3,860 respondents, 29.7% gave identical answers to Q and S, while 46.7% reported giving more coupons than they might receive.The remaining 23.6% reported the opposite; the median difference was 0 and the mean difference was 1.5 more coupons given than received.
- Reciprocation: After normalizing the difference by the maximum response, the median remained 0, with mean 0.40 and third quartile 0.67.The normalized measure more closely reflects the full reciprocity requirement but was considered more vulnerable to reporting-accuracy concerns.
S4 Measurement of Degree
The study assessed whether a one-week recall period for network degree was reasonable using reachability, coupon-distribution timing, interview-date gaps, and test-retest comparisons. Most recruitment activity occurred within a week, but time-bounded degree questions may reduce reliability.
- S4.1 Time dynamics: Within one day, the across-site-average percentage of reachable alters was 62%.Reachability was evaluated across specified time frames using reported contacts, excluding logically inconsistent responses.
- S4.1 Time dynamics: 64% of coupons were distributed in one day and 95% within seven days across sites.These self-reported distribution times support the practical relevance of a one-week time frame.
- S4.1 Time dynamics: 79% of recruiter-recruit interview pairs occurred within a week of the recruiter’s interview.This measure did not rely on respondent reports, although interview-site processing capacity may also influence interview-date differences.
- Conclusion: The three time-dynamics results suggest that restricting social-network recall to contacts seen within the previous week was reasonable in this study.The authors also note that the concentration of coupon distribution within a few days could support a two- or three-day recall period.
- S4.2 Test-retest reliability: The median test-retest difference for degree question G was 0, with 25th and 75th percentiles of -3 and 3.Results were similar across survey sites and study populations, with boxplots omitting points outside the whiskers.
- S4.2 Test-retest reliability: Non-time-bounded degree questions had higher test-retest correlations than question G, but only slightly higher.Question G’s seven-day frame can introduce week-to-week variation even when respondents answer accurately.
S5 Testing Recruitment Bias on Employment Status
The study tested employment-status recruitment patterns against non-parametric null distributions based on simple random sampling from reported eligible contacts. Very small p-values suggested recruitment bias, though inconsistent records raised data-quality concerns.
- Assumption: RDS inference commonly assumes recruits are selected randomly from each recruiter’s contacts, making recruitment-bias diagnostics important.The paper recommends Recruitment Bias Plots as an initial way to assess whether recruitment bias may be a concern.
- Method: The tests simulated simple random samples from each recruiter’s reported eligible alters and compared unweighted employment counts with null distributions.Separate procedures addressed coupon passing, coupon returns, and overall recruitment.
- Results: Very small p-values suggested that the reported recruitment patterns were very unlikely under recruitment without bias.The tests examined employment status at coupon passing, coupon return, and overall recruitment levels.
- Caveat: Many logically inconsistent records complicate interpretation because reported employed alters or recruits sometimes exceeded the corresponding available contacts or coupons.The authors identify poor data quality, possibly including desirability bias in employment-status reporting, as one explanation for extreme findings.
- Related approaches: Earlier approaches compared implied and observed group proportions through correlations or t-tests, whereas this analysis aimed to test whether compositions were the same.The paper characterizes correlation alone as inadequate for that equality question.