Source-linked AI summary

Wiki surveys: Open and quantifiable social data collection

Matthew J. Salganik, Karen E. C. Levy

arXiv:1202.0500v2stat.APcs.CY

TL;DR

The paper addresses the tension between quantifiable survey methods and open-ended approaches that capture unanticipated information. It proposes wiki surveys, including a pairwise form built around greedy, collaborative, and adaptive principles, and reports that two case studies yielded difficult-to-obtain insights. The authors also identify consistency and validity as areas requiring further study.

  • Problem

    Social-science data collection must balance quantification with openness to unanticipated information, which existing closed and open methods provide unevenly.

  • Method

    The paper develops greedy, collaborative, and adaptive wiki-survey principles and applies them to pairwise data collection and analysis in two case studies.

  • Results

    Two case studies show that pairwise wiki surveys can yield high-scoring user-contributed ideas and other information difficult to obtain from traditional surveys or interviews.

  • Takeaways & Limitations

    Pairwise wiki surveys can combine respondent-driven information growth with quantifiable evaluation of contributed ideas.

  • Takeaways & Limitations

    Further studies are needed to assess the consistency and validity of pairwise wiki-survey responses, especially because subjective attitudes rarely have gold-standard measures.

Abstract

from arXiv · show

In the social sciences, there is a longstanding tension between data collection methods that facilitate quantification and those that are open to unanticipated information. Advances in technology now enable new, hybrid methods that combine some of the benefits of both approaches. Drawing inspiration from online information aggregation systems like Wikipedia and from traditional survey research, we propose a new class of research instruments called wiki surveys. Just as Wikipedia evolves over time based on contributions from participants, we envision an evolving survey driven by contributions from respondents. We develop three general principles that underlie wiki surveys: they should be greedy, collaborative, and adaptive. Building on these principles, we develop methods for data collection and data analysis for one type of wiki survey, a pairwise wiki survey. Using two proof-of-concept case studies involving our free and open-source website www.allourideas.org, we show that pairwise wiki surveys can yield insights that would be difficult to obtain with other methods.

1 Introduction

Wiki surveys are proposed as hybrid research instruments that combine quantifiable data collection with openness to unanticipated respondent information. The paper develops this approach through three principles and pairwise-survey methods, arguing that case studies reveal insights traditional methods might miss.

  • Traditional surveys are easier to quantify and analyze, whereas open approaches better capture unexpected information but are less practical.The paper frames wiki surveys as a response to this longstanding trade-off.
  • Nearly 60% of respondents in one open-versus-closed question study supplied job-value answers outside the researchers’ five predefined choices.
  • Wiki surveys let respondents contribute new information that can produce quantifiable results while reducing researchers’ imposition of prior knowledge and biases.
  • Wiki surveys are presented as complements to, rather than replacements for, traditional closed and open methods.
  • The paper defines wiki surveys through three principles—greedy, collaborative, and adaptive—and develops data-collection and analysis methods for pairwise wiki surveys.
  • Two proof-of-concept case studies using All Our Ideas show that pairwise wiki surveys can yield insights difficult to obtain through other methods.

2 Wiki surveys

Wiki surveys adapt crowdsourced information-aggregation principles for survey research by allowing variable participation, respondent-generated content, and continual instrument adjustment. Their three principles—greedy, collaborative, and adaptive—address limits of fixed, researcher-authored, static surveys.

  • Wiki surveys should be greedy, collecting as much or as little information as each participant is willing to provide.This contrasts with traditional surveys’ fixed information quota for every respondent.
  • Online aggregation systems provide the model for these principles because they accommodate both heavy and light contributors, unlike traditional surveys.
  • Wiki surveys should be collaborative, allowing respondents to contribute new information that future respondents can evaluate.This extends beyond a traditional survey’s “other” box by integrating contributed information into later choice sets.
  • Wiki surveys should be adaptive, continually optimizing questions, order, or possible answers to elicit useful information as learning accumulates.

3 Pairwise Wiki Surveys

Pairwise wiki surveys let respondents both compare existing items and add new ones, while adaptive pair selection supports efficient learning. The analysis estimates respondent-item preferences in an opinion matrix and summarizes them with interpretable item scores.

  • Survey design: A pairwise wiki survey presents one question with many possible items, allowing respondents to compare two items and add new items for future respondents.The format operationalizes collaborative participation through respondent-generated additions.
  • Survey design: Pairwise comparisons support greedy, collaborative, and adaptive data collection by varying respondent effort, integrating new items, and selecting informative pairs.The instrument can present as many or as few prompts as respondents will answer, incorporate contributed items, and choose pairs to maximize learning.
  • Survey design: The survey format makes gaming difficult, forces prioritization between alternatives, and can provide an enjoyable sequence of web-based comparisons.Respondents cannot choose which pairs they see, must select one option from each pair, and may continue responding as long as they wish.
  • Data collection: 1,436 respondents supplied 31,893 responses and 464 ideas over about four months in the New York City sustainability survey.The survey began with 25 seed ideas and allowed respondents to contribute additional ideas subject to creator approval.
  • Data analysis: The analysis first estimates an opinion matrix with one respondent row and one item column, then summarizes it with scores estimating each item's chance of beating a random item for a random respondent.The matrix contains respondent-specific item values, while scores make the high-dimensional estimates more interpretable.
  • Data analysis: Estimating the opinion matrix requires modeling unequal response counts, unseen respondent-item combinations, and relative rather than absolute preferences.The authors use a Thurstone-Mosteller model with normally distributed item preferences and Bayesian inference, while noting that alternative models may improve the approach.

4 Case studies

Two case studies show that pairwise wiki surveys expanded the pool of policy ideas while surfacing highly ranked respondent contributions, including novel information and alternative framings.

  • Quantitative results: The PlaNYC and OECD surveys expanded active ideas from 25 to 269 and from 60 to 285, respectively.These increases represented a 10-fold expansion for PlaNYC and a 5-fold expansion for OECD.
  • Quantitative results: Respondent participation was highly unequal, with fat-head and long-tail distributions for both responses and contributed ideas.Restricting participation to 10 responses per respondent would have discarded approximately 75% of responses in each survey, while limiting contributions to one idea would have excluded nearly half of PlaNYC’s and about 40% of OECD’s user-contributed ideas.
  • Qualitative results: User-contributed ideas comprised 8 of PlaNYC’s and 7 of OECD’s 10 highest-scoring ideas.The case studies therefore identified highly supported ideas that respondents introduced rather than selecting from the seed lists.
  • Qualitative results: Many top-scoring respondent ideas arose because user contributions were more numerous and more variable in score than seed ideas.This combination made user-contributed ideas likely sources of extreme high scores even when their mean score was lower than that of seed ideas.
  • Qualitative results: Interviews classified high-scoring respondent contributions as novel information or alternative framings of existing ideas.Examples included connections across policy silos and unexpected wording for familiar policy goals.
  • Qualitative results: Taken together, the case studies suggest that pairwise wiki surveys can provide information difficult to gather through traditional surveys or interviews.The distinctive information concerned both the content of ideas and the language used to frame them.

5 Discussion

The paper presents wiki surveys as a new class of data-collection instruments and reports proof-of-concept evidence from pairwise surveys. It also identifies substantial work needed to assess their measurement properties, statistical methods, and integration with probabilistic sampling.

  • Contribution: Wiki surveys combine traditional survey research with online information aggregation to create greedy, collaborative, and adaptive data-collection instruments.Pairwise wiki surveys operationalize these principles through respondent contributions and adaptive pair selection.
  • Contribution: Two case studies show that pairwise wiki surveys can enable data collection that would be difficult with other methods.The cases used the All Our Ideas platform for community idea collection and prioritization in policymaking.
  • Future research: Additional studies are needed to assess the consistency and validity of pairwise wiki-survey responses.Suggested assessments include repeated-pair consistency, transitivity, discriminant validity, construct validity, and predictive validity.
  • Future research: Future work should improve pair selection and statistical models for estimating the opinion matrix while accounting for respondent participation.The paper notes that maximizing information per pair may not maximize information per respondent.
  • Sampling: The case studies could not use probabilistic sampling, but pairwise wiki surveys can be combined with probabilistic sampling designs.The authors frame pairwise wiki surveys as a new way of interacting with respondents rather than a new way of sampling.
  • Infrastructure: The authors made pairwise wiki surveys easier to develop by providing hosting, downloadable survey data, and open-source code.These steps are intended to facilitate further research and development.

Website implementation

The website implements pair selection, score estimation, and data-quality procedures with simple heuristics. These procedures support adaptive sampling and real-time scoring, while introducing computational and measurement trade-offs.

  • Implementation: The website addresses pair selection, score estimation, and data quality using relatively simple heuristic approaches.These were the three main methodological issues encountered during implementation.
  • Selection of pairs: Pair selection gives newer pairs opportunities to catch up with older pairs in response counts.Uniform sampling would otherwise produce more precise estimates for seed items than for user-contributed items.
  • Selection of pairs: The draw-wise probability uses the number of votes, a weighting parameter α, a probability throttle τ, and normalizing constants.The initial settings were α = 1 and τ = 0.05, although their optimal values remain an open question.
  • Score estimation: The real-time website score uses a simpler estimator than the paper’s statistical method, and the two estimates correlated about 0.95 in both cases.The simpler estimator was required because real-time calculation made the paper’s statistical methods impractical for the website.
  • Score estimation: The website score estimates an item’s probability of beating a randomly chosen item for a randomly chosen respondent on a 0-to-100 scale.The estimator derives from a binomial model with a uniform prior and adds smoothing through wins and losses.
  • Score estimation: The simpler score estimator does not account for nested responses within respondents or the strength of schedule.The paper’s more complex model addresses these limitations but requires many hours to compute.
  • Data quality: Data-quality procedures flag some responses as invalid, while the lack of login authentication means sessions are not guaranteed to represent unique respondents.The website avoided authentication to minimize barriers that might create differential non-response.

SI 2 Data analysis

The analysis estimates respondent-by-item opinion values from pairwise votes using a hierarchical model, probit response assumptions, and Gibbs sampling. It reduces computation by separating vote-informed from unobserved parameters and simplifying the design matrix.

  • Model specification: The model estimates an opinion matrix Θ representing how much each respondent values each item from pairwise responses.It explicitly models respondent-item preference heterogeneity, but the resulting case studies involved about 375,000 parameters each.
  • Model specification: Pairwise outcomes are modeled as a function of the difference between a respondent’s appeals for the two items, using the Thurstone-Mosteller probit formulation.The cumulative standard normal maps appeal differences from −∞ to ∞ into probabilities from 0 to 1.
  • Model specification: The response data are represented as a design matrix X and outcome vector Y, with each row encoding the chosen side of a pair.X places 1 on the respondent-item appearing on the left and −1 on the respondent-item appearing on the right; Y records whether the left item was chosen.
  • Hierarchical structure: Hierarchical terms enable partial pooling across respondents and allow estimation of preferences for items that particular respondents never encountered.Parameters informed by specific votes are separated from parameters inferred through hyper-parameters and other respondents’ data.
  • Computation: The posterior is estimated with Markov chain Monte Carlo using Gibbs sampling after introducing a latent continuous outcome for each observed binary response.Three parallel chains ran for 200,000 steps, with every 200th draw saved and the first half discarded as burn-in.
  • Computation: Removing columns for parameters not informed by votes made the reduced design matrix about 90% smaller in both case studies.This reduction substantially lowered the computing time and RAM required for posterior sampling.
Loading 1202.0500v2…