Source-linked AI summary

The Online Laboratory: Conducting Experiments in a Real Labor Market

John J. Horton, David G. Rand, Richard J. Zeckhauser

arXiv:1004.2931v1cs.HC

TL;DR

The paper asks whether online labor markets can provide a practical and valid alternative to laboratory and field experiments. It analyzes their experimental advantages and limitations, replicates classic findings, and concludes that online experiments can be valid, with external validity depending on the research question and researcher judgment.

  • Problem

    Laboratory experiments remain costly and inconvenient despite the potential of online labor markets, while online methods face unresolved validity challenges.

  • Method

    The paper evaluates online labor markets as experimental platforms, replicates classic experiments, conducts a labor-supply field experiment, and examines validity threats and experimental designs.

  • Results

    The authors successfully replicate selected experimental findings quickly and inexpensively, with behavior qualitatively consistent with prior findings despite MTurk’s anonymity.

  • Takeaways & Limitations

    Online labor-market experiments can be just as valid as laboratory and field experiments, but external validity depends on the research question and requires researcher judgment.

  • Takeaways & Limitations

    The external-validity domains of online results are not predetermined and require researchers’ judgment.

Abstract

from arXiv · show

Online labor markets have great potential as platforms for conducting experiments, as they provide immediate access to a large and diverse subject pool and allow researchers to conduct randomized controlled trials. We argue that online experiments can be just as valid---both internally and externally---as laboratory and field experiments, while requiring far less money and time to design and to conduct. In this paper, we first describe the benefits of conducting experiments in online labor markets; we then use one such market to replicate three classic experiments and confirm their results. We confirm that subjects (1) reverse decisions in response to how a decision-problem is framed, (2) have pro-social preferences (value payoffs to others positively), and (3) respond to priming by altering their choices. We also conduct a labor supply field experiment in which we confirm that workers have upward sloping labor supply curves. In addition to reporting these results, we discuss the unique threats to validity in an online setting and propose methods for coping with these threats. We also discuss the external validity of results from online domains and explain why online results can have external validity equal to or even better than that of traditional methods, depending on the research question. We conclude with our views on the potential role that online experiments can play within the social sciences, and then recommend software development priorities and best practices.

1 Introduction

The paper argues that online labor markets address recruitment, payment, and internal-validity barriers while offering diverse subjects, market context, and efficient experimentation. It presents replications and discusses validity, design, ethical challenges, and best practices for online experiments.

  • Motivation and contribution: Online labor markets address the recruitment/payment and internal-validity problems that have limited online experimentation.They enable large-scale subject recruitment and individual-specific payments while supporting experimental control.
  • Motivation and contribution: They provide large, diverse subject pools and allow experiments to be conducted remotely with less inconvenience and cost.Subjects are less experiment-savvy than traditional laboratory participants, and researchers need not physically aggregate them.
  • External validity and design: Experimenter-as-employer designs reduce some artificiality associated with conventional laboratory experiments and can provide high external validity for certain economic questions.Workers perform real market tasks rather than only contrived laboratory tasks.
  • Internal validity: Online labor markets support internally valid experiments through individual-specific payments, account screening, and limits on subject communication.These market features provide control needed for causal inference.
  • Challenges and scope: Online experiments are not simply laboratory experiments conducted online and raise distinct validity, ethical, and design challenges.The paper discusses threats to causal inference, external validity, experimental designs, ethics, and software priorities.
  • Evidence and scope: The paper uses replications to show that selected online experiments can reproduce established findings quickly, cheaply, and easily.The authors present these confirmations as evidence that online experiments work in the cases selected for replication.

2 Online Labor Markets and Experimentation

Online labor markets provide large, diverse pools and experimental control at low cost, while addressing recruitment and trust problems. They also introduce validity, coordination, and participant-management challenges that require safeguards.

  • Access and infrastructure: MTurk provides an experimentation-ready venue with a robust API, flexible pricing, and interfaces that web developers can adapt to many environments.The paper reports that all experiments discussed were conducted on MTurk.
  • Validity and control: Online experiments face threats from uncertain sample composition, multiple accounts, collusion, automated workers, and ethical or payment concerns.Markets address some risks through reputation systems, screening, membership management, CAPTCHA checks, escrow, and precise payments.
  • Access and infrastructure: Online labor markets rapidly recruit very large, diverse samples spanning skill levels, countries, and experience with economic games.Using workers from less-developed countries can create relatively high-stakes games at lower cost.
  • Access and infrastructure: Existing markets solve the chicken-and-egg problem by already connecting experimenters with workers, while offering low costs and speedy subject acquisition.Building a new experimentation site would otherwise require attracting both users and experiments simultaneously.
  • Validity and control: Workers’ ordinary labor-market context and limited awareness of experimentation can reduce artificiality and experimenter or John Henry effects.Workers recruited from these markets are already making consequential economic decisions and may interpret tasks economically.
  • Validity and control: Online treatment comparisons can be more accurate when subjects do not know others’ treatments and experimenters need no agents to equalize outcomes.The paper links this design to minimal demoralization and untainted treatment-control comparisons.
  • Trust and safeguards: Trust is central because valid experiments require subjects to believe that rules, payments, and stated information will be honored; online markets build mechanisms to foster it.The paper describes reputation systems, participant screening, active suspension of bad actors, and escrow as trust-supporting measures.

3 Experiments in the Online Laboratory

The authors conducted three online replications of classic behavioral findings and a labor-supply field experiment on MTurk. Results reproduced framing, other-regarding preferences, priming effects, and upward-sloping labor supply, quickly and consistently with traditional laboratories.

  • 3.2 Replication: Framing: The framing manipulation reversed preferences between equivalent gain and loss problems on MTurk.In gains, 69% chose Program A; in losses, 59% chose Program B, with p < 0.001.
  • 3.3 Replication: Social preferences: MTurk subjects cooperated significantly more than zero in a one-shot prisoner’s dilemma, replicating other-regarding preferences.Cooperation was 55% (p < 0.001).
  • 3.4 Replication: Priming: A religious prime increased prisoner’s-dilemma cooperation among believers but not non-believers.The interaction between prime and belief was positive and significant (coefficient 2.15, p = 0.001).
  • 3.5 Replication: Labor supply on the extensive margin: Workers were more likely to accept additional transcription work when offered higher wages, confirming upward-sloping labor supply.The extensive-margin response was positive across offer amounts.
  • 3 Experiments in the Online Laboratory: The replications were completed on MTurk in less than 48 hours and produced behavior qualitatively consistent with standard laboratory findings.The authors present online labor markets as a fast, inexpensive setting for experimental research.

4 Internal Validity

Online experiments face threats to internal validity from repeated participation, nonrandom attrition, timing, and participant interaction. The authors describe blocking, monitoring, market safeguards, and attrition controls to mitigate these threats.

  • Threats to internal validity: Online participants can enter over time, drop out easily, or interact, creating threats to assignment, attrition, and independence.These problems can jeopardize internal validity when treatment-related or communication-related selection occurs.
  • Multiple accounts and plays: Repeated participation can produce biased treatment effects and underestimated standard errors because repeat participants experience an unusual treatment without an obvious control.Multiple accounts are more possible online, although the authors argue they are usually a negligible threat.
  • Multiple accounts and plays: Remote, diverse worker pools reduce communication, while reputation systems make detectable dishonest play costly.Market operators also use technical and contractual approaches to limit multiple accounts.
  • Assignment and timing: Stratifying subjects by arrival time pairs participants closely and prevents arrival-related characteristics from biasing treatment-group composition.Blocking designs can likewise balance key nuisance factors before treatment assignment.
  • Coping with attrition: Selective attrition is especially acute online because participants can inspect treatments and quit with little time investment.The authors recommend collecting data on all arrivals and either testing attrition patterns or raising the cost of attrition; suitable tasks can drive attrition to zero.

5 External Validity

External validity depends on both sample representativeness and experimental realism, and neither is absolute across research questions. The authors argue that online experiments are particularly useful for identifying changes and mechanisms, while acknowledging selection and context boundaries.

  • Representativeness and realism: External validity has two dimensions: whether the sample represents the target population and whether the experimental setting resembles real decisions.Experiments typically simplify decisions relative to real life, which involves higher stakes, complexity, and decision time.
  • Representativeness and realism: Online workers are selective relative to the general population, although the authors describe them as less selective than typical physical-laboratory student samples.Self-selection remains unavoidable even when observable demographics resemble a population of interest.
  • Research-question dependence: The research question determines which sample and context are appropriate: general theories can use broad samples, whereas subgroup or specialized-context theories require closer matching.External validity is therefore a matter of degree rather than an absolute property.
  • Changes versus levels: Experiments are more reliable for studying changes and causal mechanisms than for estimating population levels, where representative samples are essential.Online laboratories gain an advantage for iterative hypothesis generation because they recruit subjects quickly and cheaply.
  • Interpreting domain differences: Systematic online–laboratory differences can create puzzles and complement conventional experiments rather than automatically invalidate online evidence.The authors report good agreement between their results and traditional methods while recognizing that measurable differences may remain.

6 Experimental Designs

Online labor markets support several experimental designs, including surveys, real-effort tasks, games, and employer-style field experiments. Their flexibility and high-frequency data advantages coexist with limits on physical tasks, physiological measurement, and synchronous interaction.

  • Design advantages: Certain research designs work online as well as or better than offline approaches, especially experimenter-as-employer natural field experiments.Workers may experience the interaction as an ordinary task and remain unaware they are participating in an experiment.
  • Task design: Online experiments recruit workers from the market and randomly assign them to groups to study incentives, team composition, framing, or payment effects.Depending on institutional review requirements, workers may not need to be told that the task is an experiment.
  • Constraints and solutions: Physical activities and physiological responses cannot be conducted or recorded online, and asynchronous worker arrival complicates subject interactions.Strategy-method designs and delayed matching offer ways to study interactive situations, while synchronous play requires additional software or coordination.
  • Design advantages: Online tasks can generate high-frequency records of workers’ actions and the context surrounding their choices.These data can reveal unexpected insights into behavior over time.
  • Task design: Text transcription and dot-guessing provide culturally neutral tasks with objective or heterogeneous quality measures and easily generated instances.The dot-guessing game has a known correct answer, while subjects may differ in estimation quality and cannot simply search online for the answer.
  • Constraints and solutions: The strategy method can produce more data points than hot interactive play by collecting responses to several hypothetical offers.Hot play remains desirable for matching the laboratory experience exactly, motivating development of web-based game platforms.

7 Ethics and Community

Online laboratories create opportunities for faster, more collaborative research but also introduce ethical and credibility risks. The paper recommends transparent sharing, replication-friendly practices, stronger tools, and preservation of the no-deception norm.

  • Ethical implications: Online experiments can be run cheaply and frequently, but researchers may obtain spuriously significant results by burying negative findings.The paper also warns that researchers may retrospectively rationalize promising pilots or experiments.
  • Ethical implications: Reduced assistance from others can weaken procedural critique, while missing laboratory logs and technicians may increase opportunities to cheat.Professional norms and easy replication are presented as mechanisms for increasing contestability and discouraging misconduct.
  • Community norms: Researchers should share machine-readable experimental materials, detailed setup instructions, raw datasets, and code that transforms raw data.Programmatic cleaning and reshaping make replication failures and consequential processing choices easier to identify.
  • Community norms: Replication-friendly sharing can reduce duplicated programming effort and make published results contestable.Open-source tools can let researchers download and redeploy working survey or experiment materials.
  • Ethical implications: Deception is especially problematic in online labor markets because distrust can pollute a shared worker resource and damage experimenter reputations.The paper argues that online experiments strengthen the case for maintaining a no-deception policy.
  • Software priorities: Better-documented software is a near-term priority, including tools for games requiring simultaneous participation and an Internet-based variant of zTree.The paper also recommends leveraging existing open-source survey infrastructure where possible.

8 Conclusion

The paper concludes that online labor-market experiments can be valid, inexpensive, and faster than traditional experiments, while external validity depends on the research question. It recommends norms and software investments to strengthen their usefulness and credibility.

  • Conclusion: Online labor-market experiments can be as valid as traditional physical laboratory experiments while reducing cost and inconvenience.The authors support this conclusion through replications conducted on MTurk.
  • Future directions: As online labor markets mature, researchers may conduct experiments across additional domains and recruit panels for experiment series.The conclusion notes that other markets may offer easier access to information about workers than MTurk.
  • Conclusion: The paper reports that online experiments can replicate established findings quickly and inexpensively.The conclusion presents replication as evidence that this approach is practically feasible.
  • Conclusion: External validity is theory-dependent rather than domain-dependent, so researchers must judge where online results will transfer.The paper does not treat online or offline status alone as determining external validity.
  • Conclusion: The authors propose norms and practices intended to enhance the usefulness and credibility of online experimentation.They describe the future of the online laboratory as promising for social-science research.
Loading 1004.2931v1…