Source-linked AI summary
Incentivizing High Quality Crowdwork
Chien-Ju Ho, Aleksandrs Slivkins, Siddharth Suri, Jennifer Wortman Vaughan
TL;DR
The paper asks when financial incentives improve crowdwork quality and investigates the question through randomized Mechanical Turk experiments and a worker-behavior model. It finds that performance-based payments work especially when tasks are effort-responsive, bonuses are sufficiently salient, and workers’ payment beliefs shape effort.
Problem
Crowdsourcing quality varies widely, while prior evidence on financial incentives and performance-based payments has been mixed.
Method
The authors run randomized Mechanical Turk experiments and propose a principal-agent model incorporating workers’ subjective beliefs about being paid.
Results
Performance-based payments improve quality when tasks are effort-responsive, bonuses are sufficiently large, and workers’ subjective payment beliefs affect effort.
Takeaways & Limitations
Requesters can pilot-test effort responsiveness using the relationship between time spent and quality when deciding whether to use performance-based payments.
Takeaways & Limitations
The handwriting-recognition experiment had little room to test thresholds because average performance was already very high.
Abstract
from arXiv · showhide
We study the causal effects of financial incentives on the quality of crowdwork. We focus on performance-based payments (PBPs), bonus payments awarded to workers for producing high quality work. We design and run randomized behavioral experiments on the popular crowdsourcing platform Amazon Mechanical Turk with the goal of understanding when, where, and why PBPs help, identifying properties of the payment, payment structure, and the task itself that make them most effective. We provide examples of tasks for which PBPs do improve quality. For such tasks, the effectiveness of PBPs is not too sensitive to the threshold for quality required to receive the bonus, while the magnitude of the bonus must be large enough to make the reward salient. We also present examples of tasks for which PBPs do not improve quality. Our results suggest that for PBPs to improve quality, the task must be effort-responsive: the task must allow workers to produce higher quality work by exerting more effort. We also give a simple method to determine if a task is effort-responsive a priori. Furthermore, our experiments suggest that all payments on Mechanical Turk are, to some degree, implicitly performance-based in that workers believe their work may be rejected if their performance is sufficiently poor. Finally, we propose a new model of worker behavior that extends the standard principal-agent model from economics to include a worker's subjective beliefs about his likelihood of being paid, and show that the predictions of this model are in line with our experimental findings. This model may be useful as a foundation for theoretical studies of incentives in crowdsourcing markets.
1. INTRODUCTION
The paper addresses mixed evidence on improving crowdwork quality by studying when and why performance-based payments work. It combines randomized Mechanical Turk experiments with a worker-behavior model focused on task responsiveness, payment design, and subjective beliefs.
- Crowdsourcing quality varies widely across tasks and workers, while existing quality-improvement techniques have had mixed success.
- The study examines performance-based payments to identify how payment amount, structure, and task properties affect crowdwork quality.
- In proofreading, bonuses improved quality across a wide range of thresholds when the bonus was sufficiently large.
- Higher base payments and performance-based payments affected quality differently, while payment timing did not explain the observed improvement.
- Performance-based payments helped tasks where additional effort was associated with higher quality, motivating an effort-responsiveness criterion for requesters.
- The paper extends the principal-agent model with workers’ subjective beliefs about payment likelihood to explain its experimental findings.
2. PRELIMINARIES
The experiments were conducted on Amazon Mechanical Turk, where requesters post paid HITs and workers choose which tasks to accept. The study uses threshold-based payments composed of a base payment, bonus, and quality threshold.
- Amazon Mechanical Turk lets requesters post HITs with offered base payments that workers can browse and accept.
- HIT qualifications restrict eligibility and can support geographic limits, repeat-participation limits, or random assignment.
- A threshold-based performance payment consists of a base payment, a bonus payment, and a threshold determining bonus eligibility.
- The experiments used a fixed $0.50 base payment, and all experiments received Microsoft Research IRB approval.
3. DOES PBP WORK?
The first experiment tests whether performance-based payments improve proofreading quality and whether fixed payments are implicitly performance-based. In a randomized six-treatment design, workers proofread articles with injected spelling errors under different bonus and acceptance conditions.
- Experiment Design: Workers proofread 500–700-word articles containing 20 randomly inserted typos and reported each typo’s location, misspelling, and correction.
- Experiment Design: Proofreading was chosen because more careful reading or repeated passes can increase the number of errors found, making quality effort-responsive.
- Results: 1,000 unique workers were assigned across six treatments, with the number of participants completing treatments not significantly different (p = 0.38).
- Results: 1.3 more typos were found with the performance-based bonus than with no bonus under non-guaranteed payment (p = 0.042).
- Results: Guaranteed base payments reduced typos found by 1.5 in the no-bonus condition and 1.3 in the unconditional-bonus condition, suggesting implicit performance beliefs.
- Results: 1.3 more typos were also found with an unconditional bonus than with no bonus under non-guaranteed payment (p = 0.036).
- Results: Performance-based payments achieved comparable quality to unconditional bonuses while costing $0.97 versus $1.50 per worker under non-guaranteed payment.
4. WHEN DOES PBP WORK?
PBPs improve proofreading quality across a broad range of bonus thresholds when the bonus is sufficiently salient, but very small bonuses provide little benefit. Thresholds that are too high may slightly reduce performance.
- Bonus Thresholds: PBPs improved proofreading quality at 25% and 75% thresholds, with workers finding 1.1 and 1.2 more typos than control, respectively.The 75% comparison was significant (p = 0.049), while the 25% comparison was not (p = 0.082).
- Bonus Thresholds: A 100% threshold did not significantly improve quality over control or the 75% threshold and slightly reduced average performance.Workers may give up when they view the bonus as unattainable.
- Bonus Thresholds: Equivalent nominal thresholds produced different outcomes: the 25% condition yielded 1.9 more typos than the five-typo condition.The difference was significant (p < .001).
- Bonus Amounts: An extra $1 of bonus produced 1.4 additional typos on average, while larger bonuses showed diminishing returns.The regression estimate was statistically significant (p = 0.002).
- Bonus Amounts: PBPs improved quality over a wide range of thresholds when the bonus was high enough to make the extra reward salient.This helps explain why very small bonuses may fail to improve quality.
5. WHY DOES PBP WORK?
The experiments indicate that higher payment itself can improve quality, while PBPs can produce still higher quality at the same total payment. The absence of an unexpected-bonus effect points toward incentive-responsive effort rather than reciprocity as the main explanation.
- Experimental Design: The experiment randomly assigned 800 workers across four treatments, with 200 workers per treatment.The completion counts did not differ significantly across treatments (p = 0.90).
- Payment Schemes: High Base and Unexpected Bonus treatments both produced more correct answers than Low Base, with p = 0.030 and p = 0.047, respectively.High Base and Unexpected Bonus did not differ significantly from each other.
- Payment Schemes: The results provide no evidence of an unexpected bonus effect in this experiment.High Base and Unexpected Bonus had the same total payment but differed in how and when payment was described.
- Payment Schemes: The PBP treatment outperformed all other treatments, with p < 0.005.Workers knew about the bonus before accepting the HIT.
6. WHERE DOES PBP WORK?
PBPs improved quality on tasks where additional time was associated with better work, but not on tasks lacking that relationship. The authors therefore use effort-responsiveness as a practical criterion for deciding when PBPs may help.
- Effort-Responsive Tasks: For proofreading, each additional minute correlated with finding 0.42 more typos; for spot-the-difference, it correlated with 0.17 more correct answers.Both relationships were statistically significant (p < 0.001).
- Scope: The comparison with prior work is bounded by different settings: the cited prior study used three-hour oDesk tasks, whereas these MTurk tasks averaged 8.8 minutes.The authors explicitly note that the experimental settings are not the same.
- Handwriting Recognition: Handwriting recognition was not effort-responsive, and PBPs did not significantly improve accuracy over control.Average control accuracy was already 95.2%, leaving limited room for improvement.
- Audio Transcription: Audio transcription was not effort-responsive, and none of the three PBP treatments significantly outperformed control.The absence of a PBP effect was not fully explained by a ceiling effect because control accuracy averaged 75.4%.
- Practical Recommendation: The authors’ noncausal evidence supports effort-responsiveness as an important reason PBPs help on some tasks but not others.They propose piloting a task with fixed payment and relating time spent to work quality.
7. A THEORY OF WORKER INCENTIVES
The worker model extends principal-agent reasoning by incorporating subjective beliefs about payment and acceptance, explaining when performance-based payments improve quality. It predicts that PBPs help when extra effort can raise quality at reasonable cost and when the bonus and threshold make that effort worthwhile.
- Worker model: The model defines worker utility as perceived expected payment minus the perceived cost of producing a chosen quality level.Under PBPs, expected payment includes base and bonus payments weighted by workers’ perceived probabilities of receiving them.
- Scope of explanation: The model explains empirical observations that the standard principal-agent model cannot explain, but its quality predictions rely on uncertainty about receiving payment.Without payment uncertainty, increasing guaranteed pay would not increase quality in the standard model.
- Model consequences: Subjective beliefs about payment and acceptance criteria explain why implicit performance-based payments and higher payments can increase quality.The model predicts quality increases relative to guaranteed or standard payments when workers perceive payment receipt as quality-dependent.
- Conditions for effectiveness: PBPs help when additional effort can raise quality at reasonable cost, while the bonus must be sufficiently large and the quality threshold must meaningfully distinguish high from low quality.If producing high quality already costs little, or if the bonus is too small or the threshold too low, PBPs may not improve quality.
- Threshold design: Equation 3 offers a simple way to reason about whether changing the perceived bonus threshold can improve incentives.Workers’ perceptions of the threshold may differ even when alternative bonus rules award payments at roughly the same objective frequency.
8. CONCLUSIONS AND DISCUSSION
The experiments show that PBPs improve quality for some tasks but not others, with task effort-responsiveness helping explain this difference. The paper also identifies payment-design conditions, implicit performance incentives, and a worker model that organizes these findings.
- Overall findings: PBPs can improve submitted-work quality for some tasks but are not likely to do so for others.The conclusions identify task characteristics as central to whether PBPs work.
- Task selection: A task’s effort-responsiveness is a potential reason PBPs succeed or fail, measurable by whether more time spent is associated with sufficiently higher quality.The paper proposes a pilot experiment using the correlation between time spent and quality to assess this property before adopting PBPs.
- Implicit incentives: Workers may treat fixed payments as implicitly performance-based because they hold subjective beliefs about the quality required for acceptance.Requesters may also use relative or inaccessible gold-standard thresholds to influence perceived bonus likelihood.
- Payment design: PBPs require a sufficiently large bonus, exhibit diminishing returns as the bonus increases, and can improve quality across a wide range of quality thresholds.These findings help explain why prior studies reported inconsistent effects of PBPs.
- Theory: The proposed worker-behavior model captures the reported payment and task effects and provides a foundation for further theoretical work on crowdsourcing markets.It is intended as a more realistic way to analyze payment schemes than the standard principal-agent model alone.