Source-linked AI summary

Crowdsourced Judgement Elicitation with Endogenous Proficiency

Anirban Dasgupta, Arpita Ghosh

arXiv:1303.0799v1cs.GT

TL;DR

Crowdsourced evaluation requires incentives for both effort and truthful reporting when ground truth and effort are unobservable. The paper proposes a multi-task binary elicitation mechanism that penalizes low-effort agreement, making full-effort truthful reporting the highest-payoff equilibrium. Its guarantees require only minimal prior bounds and own-evaluation reports, without requiring diverging task-level participation.

  • Problem

    Crowdsourced judgement elicitation must induce effort and truthful reporting when agents’ proficiencies depend on strategically chosen effort and ground truth is unobservable.

  • Method

    The paper uses multiple tasks and ratings to reward agreement with a reference rater while subtracting agreement expected from reporting statistics, thereby penalizing blind agreement.

  • Results

    Full-effort truthful reporting is a Nash equilibrium with maximum payoff to all agents, including under heterogeneous proficiencies, mixed strategies, and task-specific strategies.

  • Takeaways & Limitations

    The mechanism provides binary information elicitation with minimal prior bounds, own-evaluation reports, and no diverging-report requirement for its incentive properties.

  • Takeaways & Limitations

    The mechanism can have other equilibria, including all agents using low effort and receiving zero reward.

Abstract

from arXiv · show

Crowdsourcing is now widely used to replace judgement by an expert authority with an aggregate evaluation from a number of non-experts, in applications ranging from rating and categorizing online content to evaluation of student assignments in massively open online courses via peer grading. A key issue in these settings, where direct monitoring is infeasible, is incentivizing agents in the `crowd' to put in effort to make good evaluations, as well as to truthfully report their evaluations. This leads to a new family of information elicitation problems with unobservable ground truth, where an agent's proficiency- the probability with which she correctly evaluates the underlying ground truth- is endogenously determined by her strategic choice of how much effort to put into the task. Our main contribution is a simple, new, mechanism for binary information elicitation for multiple tasks when agents have endogenous proficiencies, with the following properties: (i) Exerting maximum effort followed by truthful reporting of observations is a Nash equilibrium. (ii) This is the equilibrium with maximum payoff to all agents, even when agents have different maximum proficiencies, can use mixed strategies, and can choose a different strategy for each of their tasks. Our information elicitation mechanism requires only minimal bounds on the priors, asks agents to only report their own evaluations, and does not require any conditions on a diverging number of agent reports per task to achieve its incentive properties. The main idea behind our mechanism is to use the presence of multiple tasks and ratings to identify and penalize low-effort agreement: the mechanism rewards agents for agreeing with a `reference' rater on a task but also penalizes for blind agreement by subtracting out a statistic term designed so that agents obtain reward only when they put effort into their observations.

1 Introduction

The paper addresses crowdsourced evaluation when agents’ effort affects proficiency and ground truth is unobservable. It introduces a binary, multi-task mechanism that incentivizes full-effort truthful reporting and penalizes low-effort agreement.

  • Motivation: Crowdsourced judgement elicitation must incentivize both truthful reporting and effort when agents form evaluations as part of their tasks.Unlike settings with pre-formed opinions, evaluation accuracy depends on strategically chosen effort.
  • Contribution: The paper introduces a simple mechanism for binary information elicitation across multiple tasks with endogenously determined agent proficiencies.The model focuses on binary judgements such as identifying adult content or evaluating correctness.
  • Incentive properties: Maximum effort followed by truthful reporting is a Nash equilibrium and gives all agents the maximum payoff across equilibria.This remains true with heterogeneous maximum proficiencies, mixed strategies, and task-specific strategies.
  • Incentive properties: If each task may have a trusted truthful rater with proficiency above half, full-effort truthful reporting is essentially the only equilibrium.The mechanism need not know which agents are trusted.
  • Scope and novelty: The mechanism requires only minimal prior bounds, uses agents’ own evaluations, and does not require a diverging number of reports per task.The paper contrasts these guarantees with prior information-elicitation mechanisms.
  • Mechanism: Multiple tasks let the mechanism distinguish high-effort agreement from blind agreement by subtracting agreement expected from agents’ reporting statistics.Agents are rewarded for agreement with a reference rater only after low-effort agreement is penalized through the statistic term.

2 Model

The model studies binary judgement elicitation when agents choose both effort and reporting, with effort determining evaluation proficiency. It seeks mechanisms making full effort and truthful reporting an equilibrium with maximum utility.

  • Tasks and agents: Agents evaluate binary-valued task qualities that remain unknown to the system.Tasks have underlying types H or L, while agents report binary evaluations based on or independently of their observations.
  • Proficiency: An agent’s proficiency is the probability of correctly evaluating a task’s true type and increases with effort.Effort is binary: zero effort costs nothing and yields random guessing, while full effort costs c_ij(1) and yields proficiency p_i.
  • Proficiency: Full-effort proficiency may differ across agents, need not be known to the center, and is assumed to be at least 1/2.The model also allows proficiency to depend on whether the object is H or L.
  • Strategies and equilibrium: Agents strategically choose an effort level and reporting function for each task to maximize rewards minus evaluation costs.Strategies can vary across an agent’s D tasks, and equilibrium expectations include evaluation and mechanism randomness.
  • Design objective: The target is for full effort followed by truthful reporting on all tasks to be an equilibrium with maximum utility, without solving the separate aggregation problem.The mechanism elicits high-quality judgements for later aggregation rather than determining the final estimate of each task’s true quality.

3 Mechanism

The mechanism compares each report with a reference rater while using reports from additional non-overlapping tasks to estimate and subtract expected blind agreement. This preserves rewards for informative agreement and removes rewards from effort-free reporting patterns.

  • Mechanism construction: M_d rewards agreement with a reference rater only after subtracting a statistic based on other reports’ empirical frequencies.The statistic uses d other reports from each of the agent and reference rater.
  • Mechanism construction: The mechanism uses d non-overlapping tasks other than j for the agent and reference rater to compute B_ij.The sets S_ij and S_rj(i)j each contain d tasks and are disjoint.
  • Mechanism inputs: M_d requires only agents’ reports and does not require the center to know their maximum proficiencies.Reference raters and the auxiliary task sets completely specify the mechanism.
  • Reward computation: The final reward aggregates task-level rewards and scales them by the non-negative parameter β.The supplied definition specifies the final reward as βR_i.
  • Reward components: The agreement term A_ij equals 1 when the agent and reference rater report the same binary value.This includes agreement on either H or L.
  • Reward components: Subtracting B_ij penalizes blind agreement by removing agreement expected from independent reports with the agents’ empirical reporting frequencies.If all agents always report H or use independent biased coin tosses, the net reward is 0.
  • Mechanism variants: Two natural variants set d=D−1 or d=1, respectively, determining how many auxiliary reports enter the statistic term.M_D−1 excludes the common task from the reports used for the statistic, while M_1 uses one report.

4 Analyzing M

The analysis shows that full-effort truthful reporting is an equilibrium and, under apriori-equivalent tasks, gives each agent maximum reward. The mechanism also characterizes low-effort and alternative equilibria, including conditions that make deviations strictly unprofitable.

  • Preliminaries: The parameter d has no clear effect on mechanism behavior under risk neutrality, leaving its role as an open question.The mechanism’s payments can nevertheless be scaled appropriately to preserve the full-effort truthful equilibrium for any effort costs.
  • Equilibrium analysis: Full-effort truthful reporting is an equilibrium when agents have proficiency above one-half and reference-rater conditions are satisfied.The theorem requires equal task contribution d_ij=d and equal expected reference-rater proficiency across an agent’s tasks.
  • Equilibrium analysis: Any deviation on one task or a subset of tasks strictly decreases an agent’s expected reward when d_ij=d.The total reward decomposes across task-specific terms, so the result applies even when strategies differ by task.
  • Equilibrium analysis: The mechanism can also admit other equilibria, including zero-reward random-reporting equilibria and a full-effort report-inversion equilibrium.The inverted-report equilibrium achieves the same maximum expected reward but is described as unnatural and risky, while random reporting yields zero reward.
  • Equilibrium analysis: The full-effort truthful equilibrium maximizes each agent’s reward when tasks are apriori equivalent.The reward component associated with each reference rater is maximized when both agents exert full effort and report truthfully.
  • Equilibrium analysis: Mixed equilibria involving low effort collapse to the all-low-effort equilibrium when every agent can serve as a reference rater with nonzero probability.Any agent observing a sufficiently proficient reference rater using full effort has a strict profitable deviation to full-effort truthful reporting.

5 Creating the Task Assignment

The section gives a task-assignment algorithm that makes suitable reference raters feasible under mild conditions. With m ≥ D^2, every agent receives D tasks, every task has T agents, and each reference rater shares only the target task.

  • Assignment algorithm: The construction randomly permutes agents, partitions tasks into D-sized blocks, and partitions agents into T blocks.Agents are assigned according to their block positions and task-block indices.
  • Assignment algorithm: Agents in the first block receive all tasks from their corresponding task-block, while later blocks receive shifted sets of D tasks.The shifts distribute task assignments across blocks and complete agent capacities.
  • Reference raters: Reference raters are chosen from the first agent block for later-block agents and from another worker on the task for first-block agents.This choice is designed to ensure that the reference rater and evaluated agent have only the target task in common.
  • Feasibility: Under the construction, each agent-task pair has a reference rater whose assignment intersects the agent’s assignment only at that task.This satisfies the condition required for M_D−1.
  • Feasibility: m ≥ D^2 makes the assignment feasible, with every agent assigned exactly D tasks and every task assigned to T agents.The lemma also establishes the required reference-rater conditions and an expectation condition induced by the initial random permutation.

6 Discussion

The paper presents a mechanism for crowdsourced binary information elicitation with effort-dependent proficiency. It identifies and penalizes low-effort agreement, making maximum-effort truthful reporting the highest-payoff Nash equilibrium while leaving broader outcome spaces and task heterogeneity for future work.

  • Discussion: The mechanism uses multiple tasks to identify and penalize low-effort agreement, incentivizing effort when task types are binary.It is designed for settings where agents’ proficiencies depend strategically on their effort.
  • Discussion: Maximum effort followed by truthful reporting is the Nash equilibrium with maximum payoff to all agents, including mixed-strategy equilibria.The result covers the paper’s endogenous-proficiency setting and allows agents to have different maximum proficiencies.
  • Discussion: The mechanism requires only agents’ own evaluations and no diverging number of reports per task to achieve its incentive properties.The authors position it as a starting point for image labeling, tagging, and peer grading.
  • Limitations and future work: The model assumes binary outcomes, binary effort, homogeneous tasks, and common costs and maximum proficiencies across tasks.Extending the framework to richer outcomes, convex effort costs, and heterogeneous tasks is identified as future work.
Loading 1303.0799v1…