Source-linked AI summary
Breaking Monotony with Meaning: Motivation in Crowdsourcing Markets
Dana Chandler, Adam Kapelner
TL;DR
Prior evidence on meaningful work had generally come from ethnographies, observational studies, or laboratory experiments. This paper uses a natural field experiment in MTurk to manipulate task framing and finds that meaningfulness changes participation and how workers allocate effort between quantity and quality.
Problem
Prior studies of meaningful work generally relied on ethnographies, observational studies, or laboratory experiments rather than a natural field experiment in a real labor market.
Method
The study recruited workers from the United States and India on MTurk, experimentally varied task framing, and measured participation, output quantity, and labeling quality.
Results
Meaningful framing increased participation and quantity without significantly increasing quality, while shredded framing decreased quality without changing quantity.
Takeaways & Limitations
Meaning appears to moderate how workers substitute effort between task quantity and task quality in short-term labor markets.
Takeaways & Limitations
The results match short-term labor environments like MTurk but do not easily generalize to longer-term employment relationships.
Abstract
from arXiv · showhide
We conduct the first natural field experiment to explore the relationship between the "meaningfulness" of a task and worker effort. We employed about 2,500 workers from Amazon's Mechanical Turk (MTurk), an online labor market, to label medical images. Although given an identical task, we experimentally manipulated how the task was framed. Subjects in the meaningful treatment were told that they were labeling tumor cells in order to assist medical researchers, subjects in the zero-context condition (the control group) were not told the purpose of the task, and, in stark contrast, subjects in the shredded treatment were not given context and were additionally told that their work would be discarded. We found that when a task was framed more meaningfully, workers were more likely to participate. We also found that the meaningful treatment increased the quantity of output (with an insignificant change in quality) while the shredded treatment decreased the quality of output (with no change in quantity). We believe these results will generalize to other short-term labor markets. Our study also discusses MTurk as an exciting platform for running natural field experiments in economics.
1. Introduction
The paper tests whether framing identical crowdsourcing work as meaningful changes participation, output, quality, and compensation. In a natural field experiment on MTurk, meaningful framing increased participation and quantity, while denying meaning reduced quality without reducing quantity.
- Motivation and research gap: Prior evidence on meaningful work largely came from ethnographies, observational studies, and laboratory experiments.
- Research question: The experiment asks whether employers can alter perceived task meaningfulness to induce more or higher-quality work at lower wages.
- Experimental approach: Identical pay, requirements, and working conditions were maintained while cues changed workers’ perceived meaningfulness.
- Participation: 80.6% of meaningful-treatment workers labeled at least one image, compared with 76.2% in zero context and 72.3% in shredded.
- Quantity: Meaningful framing made workers approximately 23% more likely to become high-output workers, defined as labeling at least five images.
- Dimensions of effort: Meaning increased quantity without significantly increasing quality, whereas shredding reduced quality without changing quantity.
- Compensation: Meaningful-treatment workers earned $1.34 per hour, 6 cents less than zero-context workers and 14 cents less than shredded-condition workers.
- Implications: The authors argue that meaningfulness matters substantially in short-term labor markets and present MTurk as a platform for natural field experiments.
2. Mechanical Turk and its potential for field experimentation
MTurk combines a large, inexpensive pool of workers with scalable, customizable, relatively anonymous one-off tasks. These features support natural field experiments, while generalization remains limited by selection concerns and the short-term setting.
- MTurk as a labor market: MTurk is a large task-based labor market where hundreds of thousands of people worldwide complete Human Intelligence Tasks for pay.
- Research uses: Researchers use MTurk both to outsource tasks and to conduct online experiments, including cross-national and replication studies.
- Operational advantages: MTurk allows researchers to recruit hundreds of subjects within days at substantially lower cost than laboratory experiments.
- Naturalistic environment: Its relative anonymity, one-off work, and limited worker-employer interaction reduce some concerns about scrutiny, reputation, and career effects.
- Generalizability: The study acknowledges possible unobservable differences between MTurk workers and broader populations despite observable demographic representativeness.
- Scope boundary: Results are expected to generalize to short-term work but not readily to longer-term employment relationships or occupational decisions.
- Experimental control: Researchers can create customized work environments without relying on private-sector partners, and MTurk experiments can be replicated by rerunning source code.
3. Experimental Design
The experiment randomized MTurk workers into meaningful, zero-context, or shredded framing conditions while holding task mechanics and pay structure constant. It measured participation, image-labeling quantity, labeling quality, and workers’ perceptions of the task.
- Subject recruitment: The task was posted as an ordinary one-time MTurk opportunity, and workers were barred from completing the experiment more than once.
- Subject recruitment: The study hired 2,471 workers: 1,318 from the US and 1,153 from India.
- Treatment assignment: Workers were randomized into meaningful, zero-context, or shredded conditions after providing demographics and passing a color-blindness test.
- Treatment manipulation: Meaningful framing thanked workers and explained that researchers needed help labeling cancerous tumor cells, while the other conditions used less meaningful descriptions.
- Treatment manipulation: The scripts shared the same wage structure and labeling mechanics, with only treatment-specific framing and terminology differing.
- Outcome measures: The study measured induced-to-work, quantity of image labelings, and quality of image labelings as its three response variables.
- Incentives: Workers received $0.10 for the first image, followed by declining piece rates to $0.02, where the rate remained constant.
- Outcome measures: Quality was the proportion of objects correctly identified within a selected pixel radius of each object’s true center.
4. Experimental Results and Discussion
Meaningful framing increased participation and quantity, while shredded framing reduced labeling quality without reducing quantity. These effects were broadly similar across the United States and India, though the shredded manipulation may not have changed perceived meaningfulness.
- Experiment and sample: 2,471 workers participated, including 1,318 from the United States and 1,153 from India.The experiment measured participation, image quantity, quality, demographics, and hourly wage.
- Effort allocation: The meaningful condition increased quantity without significantly increasing quality, whereas shredding decreased quality while quantity remained constant.The authors describe this allocation of effort between quantity and quality as a “checkerboard effect.”
- Cross-country comparison: Treatment effects did not differ significantly between United States and Indian workers, despite country-level differences in participation, output, and accuracy.Indian workers were less accurate and completed images less often overall, but country did not significantly moderate the treatment effects.
- Quantity results: Meaningful framing increased output, including 4.7% more workers labeling at least two images and nearly 23% more becoming high-output workers.High-output workers labeled five or more images; the shredded treatment had no corresponding output effect.
- Quality results: Fine quality was 7.2% lower in the shredded treatment, while meaningful framing produced no large corresponding increase.In the United States, meaningful framing increased fine quality by 3.9% with controls, but there was no effect in India.
- Caveat: Quality estimates are subject to attrition bias because quality was observed only for workers induced to label images.Attrition was 4% higher in the shredded treatment, and the authors presume those who opted out would have produced worse quality.
- Manipulation check: The shredded treatment may not have achieved its intended manipulation because it did not significantly lower post-task ratings on any measured item.Meaningful-treatment workers rated meaningfulness, purpose, enjoyment, accomplishment, and recognition higher.
5. Conclusion
The experiment shows that perceived meaning affects participation and how workers allocate effort across output quantity and quality. The authors also position MTurk as a useful setting for natural field experiments and suggest relevance to other short-term labor markets.
- The experiment is described as the first natural field experiment examining how task meaningfulness influences labor supply.
- Meaningful framing increased quantity with an insignificant quality change, while shredded framing reduced quality without changing quantity.
- The authors suggest that perceived meaning may affect substitution between task quantity and task quality.
- The findings have implications for employers using short-term labor, including temp-work and piecework, because non-pecuniary incentives matter significantly.
- The study presents MTurk as a platform for high-internal-validity natural field experiments that can avoid some laboratory external-validity problems.
Appendix A. Detailed Experimental Design
The appendix details how workers encountered, learned, and performed the image-labeling task. The experiment varied instructional meaning cues and whether completed work would be saved, while standardizing image content and recording additional labeling behavior.
- Task entry and screening: Workers first encountered the HIT on MTurk through a preview screen and then completed screening activities including colorblindness, demographic, and audio tests.
- Treatment assignment: Workers were randomized into meaningful, zero-context, or shredded treatments before viewing a treatment-specific instructional video.
- Treatment framing: The meaningful script explained that labeling tumor cells would help researchers, whereas the other scripts referred to generic objects of interest.
- Training and interface: Workers passed a quiz before labeling images using an interface with zoom controls, point creation and deletion, example cells, and thumbnail navigation.
- Task and incentives: Each image contained the same 90 cells arranged and rotated across believable backgrounds, and workers could label additional images at declining piece rates.
- Treatment framing: The shredded treatment additionally told workers that none of their points would be saved because the system was being tested, although they would still be paid.
Institutional Review Board (IRB) Requirements
The experiment required deception and delayed disclosure to preserve its naturalistic setting. The authors also identify communication among MTurk workers as a potential internal-validity problem.
- The study used deception because subjects could not be told initially that they were participating in an experiment.
- An IRB would most likely require a debrief statement explaining the experiment, its purpose, and institutional contact information.
- Debriefing had to occur after data collection because earlier disclosure could compromise the experiment.
- MTurk subjects may communicate with one another that a study is underway, creating a potential internal-validity problem.
Engineering Required
Running an MTurk experiment requires both front-end and back-end engineering, alongside experimental-design knowledge. The recommended setup includes automated task creation and worker payment.
- An experiment of this scale requires no more than two weeks of full-time work for an experienced software engineer.
- Front-end: Front-end engineering uses HTML and CSS for task rendering, JavaScript for dynamic behavior, and AJAX for client-server communication.
- Back-end: The back end needs an HTTP server, database, and server-side platform to render pages and store experiment data.
- Automation: MTurk integration benefits from experience with Amazon’s API and Linux CRON jobs.
- Automation: One recommended CRON creates fresh HITs every 15 minutes, while another automatically pays workers about every hour, including bonuses and rejections.
- The engineer should also understand experimental-design principles in addition to technical skills.