Source-linked AI summary
The Ordinal Annotation Game: How Construct Abstraction Shapes Crowdsourced Consensus
Kosmas Pinitas
TL;DR
The paper asks whether persistent ordinal annotation disagreement is random noise or structured coordination against an internalised population prior. It formalises the Ordinal Annotation Game and evaluates it in sensory and engagement experiments using the same interface and threshold. The payoff slope reverses sign across constructs: sensory annotation is consensus-dominant, whereas abstract engagement annotation is effort-limited.
Problem
Persistent inter-annotator disagreement is conventionally treated as random noise, despite annotators balancing interface-change effort against alignment with an internalised population prior.
Method
The paper models independent ordinal annotation as a repeated coordination game and compares sensory tracking with abstract engagement tracking under identical interface software and δ = 0.05.
Results
The payoff slope is positive for sensory tasks but negative for engagement under the same processing pipeline, producing consensus-dominant and effort-limited regimes respectively.
Takeaways & Limitations
Ordinal disagreement can be a structured behavioural signal, and a single majority-vote ground truth may be inappropriate when construct abstraction produces private appraisal criteria.
Takeaways & Limitations
The contrast is associational rather than causal because the corpora differ in stimulus complexity and participant population; the engagement cohort is small (N2 = 5).
Abstract
from arXiv · showhide
Inter-annotator disagreement in real-time affect annotation is widely treated as stochastic noise. We challenge this view by modelling ordinal annotation as an implicit game-theoretic coordination process against an internalised population prior under a post-hoc majority vote. We present the Ordinal Annotation Game, a conceptual scaffold in which the mapping from individual effort to collective consensus is governed by the semantic abstraction of the target construct. We evaluate it across two experiments sharing identical interface software and a uniform sensitivity threshold: a controlled sensory tracking study and an in-the-wild engagement study. Sensory annotation yields a consensus-dominant regime where active updates reinforce agreement, whereas engagement annotation inverts into an effort-limited regime where more labelling penalises consensus. The payoff slope reverses sign under identical processing, showing that ordinal disagreement is a structured behavioural phenomenon, not a discretisation artefact or random error.
I. INTRODUCTION
The paper reframes persistent inter-annotator disagreement as strategic coordination against an internalised population prior rather than random noise. It introduces the Ordinal Annotation Game and argues that construct abstraction governs whether effort aligns annotators or decouples them from consensus.
- I. INTRODUCTION: Ordinal disagreement is treated as a strategic coordination process rather than stochastic noise removed by filtering, smoothing, or voting.Annotators balance the effort of changing an interface state against alignment with an internalised population prior.
- I. INTRODUCTION: The Ordinal Annotation Game models independent annotation as a repeated game in which continuous perceptions become discrete directional actions under a shared deadband.The framework uses a uniform threshold to isolate construct abstraction as a factor in crowdsourced divergence.
- I. INTRODUCTION: Concrete perceptual attributes are predicted to produce consensus-dominant coordination, whereas abstract psychological constructs are predicted to produce effort-limited coordination.Private appraisal criteria in abstract tasks can make additional effort increasingly decouple individual judgments from group consensus.
- I. INTRODUCTION: Ordinal interfaces extend prior rank-based and relative annotation approaches intended to reduce reaction lag and cognitive friction.The paper contrasts its coordination perspective with reliability and fusion methods that assume a singular recoverable ground truth.
A. Action Space and Threshold Mapping
The annotation pipeline converts continuous user traces into three discrete directional actions using a post-hoc sensitivity deadband. Consensus is then computed from the instantaneous majority action after collection.
- A. Action Space and Threshold Mapping: A sensitivity deadband of half-width δ > 0 filters microfluctuations before continuous traces are mapped into a three-element ordered action space.The mapping is described as the designer’s post-hoc transformation from continuous perception to discrete ordinal labels.
- A. Action Space and Threshold Mapping: The mapping from continuous perception to discrete labels is associated with structural disagreement as constructs become more abstract.
- A. Action Space and Threshold Mapping: The action space consists of discrete directional labels −1, 0, and +1.
- A. Action Space and Threshold Mapping: The analytical population consensus is the instantaneous mode of annotators’ actions within each time window.This majority aggregate is computed post-hoc and is never shown to annotators during labelling.
B. Utility as a Conceptual Scaffold
The Ordinal Annotation Game defines annotation utility through coordination, effort, and directional bias. Its utility equation is conceptual: it formalises possible coordination dynamics rather than fitting or predicting affect trajectories.
- B. Utility as a Conceptual Scaffold: Annotator utility combines coordination with the expected consensus, a penalty for behavioural effort, and a penalty for systematic directional bias.The coordination term rewards alignment, while effort and bias represent distinct costs.
- B. Utility as a Conceptual Scaffold: Coordination is defined as whether an annotator’s action matches the expected consensus, while effort measures unsigned step-to-step label volatility.
- B. Utility as a Conceptual Scaffold: Bias measures net displacement from the initial action and remains distinct from cumulative effort.Frequent symmetric updates can produce high effort but low bias.
- B. Utility as a Conceptual Scaffold: The utility expression is a conceptual vehicle for explaining a stable action distribution shaped by stimulus saliency, aggregation, thresholding, and shared or unshared semantic parameters.The paper does not fit the equation or estimate its coefficients, and bias is not used in the empirical analysis.
- B. Utility as a Conceptual Scaffold: The Ordinal Annotation Game is defined as a repeated process in which independent annotators generate continuous evaluations that optimise an internal utility expectation under an aggregation rule and deadband.The resulting time-series values are treated as empirical samples from an action distribution.
C. The Unified Phase Transition Hypothesis
The paper tests whether construct abstraction reverses the relationship between annotation effort and consensus while holding the processing threshold fixed. Its hypothesis predicts positive effort–consensus coupling for concrete sensory constructs and negative coupling for abstract psychological constructs.
- C. The Unified Phase Transition Hypothesis: Effort and agreement are summarised as windowed metrics over an operational analytical interval.The formulation examines the empirical relationship between labelling effort and aggregate consensus agreement.
- C. The Unified Phase Transition Hypothesis: The payoff-frontier slope dA/dE is estimated by ordinary least squares from Ai = mEi+c, with m = dA/dE.Because effort and agreement derive from the same action stream, the raw slope sign is necessary but not sufficient without the sensory control.
- C. The Unified Phase Transition Hypothesis: For concrete sensory constructs, tightly coupled internal signals are predicted to yield a positive frontier slope, dA/dE > 0.
- C. The Unified Phase Transition Hypothesis: For abstract psychological constructs, private appraisal criteria are predicted to expand individual variance and drive a negative frontier slope, dA/dE < 0.
- C. The Unified Phase Transition Hypothesis: The proposed mechanism is that high effort tracks shared physical transitions in sensory tasks but produces uncoordinated private movements in abstract tasks.The identical mechanical co-determination of effort and agreement motivates using the sign reversal as the key comparison.
IV. EXPERIMENTAL DESIGN AND DATASETS
The paper compares two ordinal annotation experiments using identical interface software and the same physical threshold, δ = 0.05. Both use overlapping 3 s windows with 1 s steps, while annotators produce continuous PAGAN traces that are later discretised.
- Both experiments share the exact same interface software and physical threshold setting, δ = 0.05.
- Both datasets are segmented using overlapping 3 s windows with a 1 s step size.
- Annotators produce unbounded continuous traces in PAGAN using the mouse wheel, which are subsequently discretised into directional actions through the uniform deadband.
A. Experiment 1: Controlled Sensory Tracking
Experiment 1 measures low-level sensory changes through colour intensity and sound pitch, while the broader study models saliency and coordination on a shared discrete action space. The same PAGAN interfaces support sensory and gameplay annotation.
- Experiment 1: Controlled Sensory Tracking: Experiment 1 uses colour intensity and sound pitch as low-level sensory annotation tasks.
- Experiment 2: Abstract Psychological Tracking: Experiment 2 uses one-minute first-person shooter clips to annotate gameplay engagement.
- Exogenous Saliency: Saliency is defined as luminance or pitch change for sensory stimuli, but derives from a unified audiovisual pipeline for naturalistic gameplay.
- Metric Modelling: Physical saliency curves and group response entropy are computed within each 3 s window and transformed with δ = 0.05 into −1, 0, or +1.
- Metric Modelling: Cohen’s Kappa is used because the compared tracks already share the discrete alphabet {−1, 0, +1}.The coupling measures co-occurring discretised directions above chance rather than equality of the underlying constructs.
Methodological Limitations and Confounds:
The comparison cannot isolate construct abstraction causally because the experiments differ in stimulus complexity, participant population, and designed ground truth. The paper therefore frames the contrast as a strong association.
- The causal effect of construct abstraction is confounded by differences between the two corpora.
- Experiment 1 uses synthetic, low-complexity stimuli and a general cohort, whereas Experiment 2 uses complex naturalistic gameplay and expert researchers.
- The paper frames the contrast as a strong association rather than an isolated causal coupling.
V. EQUILIBRIUM ANALYSIS
The equilibrium analysis finds opposite effort–consensus regimes under the same processing framework: sensory annotation has positive payoff-frontier slopes, while gameplay engagement has strongly negative slopes. The temporal figures show these patterns persist across the observed horizons.
- Consensus-Dominant Regime: 0.63 for colour intensity and 0.57 for sound pitch are strongly positive population frontier slopes under δ = 0.05.The 95% bootstrap CIs are [0.48, 0.76] and [0.41, 0.71], respectively.
- Cross-Regime Comparison: The engagement slopes reverse sign relative to sensory annotation under the same uniform processing framework.
- Consensus-Dominant Regime: The sensory payoff–frontier slope dA/dEt remains robustly positive across the time horizon.
- Effort-Limited Regime: The engagement payoff–frontier slope dA/dEt stays close to −1 throughout the 60 s session, contrasting with the positive sensory frontier.
- Effort-Limited Regime: −0.98 for Visual, −0.99 for Auditory, and −0.98 for Audiovisual indicate near-perfect negative engagement slopes.The reported intervals are Visual [−0.99, −0.96], Auditory [−0.99, −0.97], and Audiovisual [−0.99, −0.95].
C. Ruling Out the Discretisation Artefact
Because both experiments use identical interface, aggregation, and sensitivity threshold, the differing frontier signs cannot be attributed solely to mechanical negative coupling. The sign reversal is therefore interpreted as reflecting construct abstraction rather than discretisation.
- Controlling the pipeline: δ = 0.05 is identical across both experiments, making any mechanical action-stream bias common to both conditions.Ei and Ai derive from the same action stream, so active annotators are mechanically more likely to deviate from the instantaneous majority.
- Controlling the pipeline: If the negative frontier were purely mechanical, both sensory and engagement tasks would exhibit the same sign.The sensory task functions as a built-in control because the interface, aggregation, and threshold are held constant.
- Interpreting the reversal: The identical pipeline makes construct abstraction the most parsimonious explanation despite differences in cohorts and corpora.The authors identify the sign reversal, rather than frontier magnitude, as the key result.
- Interpreting the reversal: dA/dEt ≈−1 across most engagement-session windows, with a sharper deviation near 52 s at approximately −2.1.The near-zero saliency couplings and the 52 s deviation are consistent with conservative annotators holding steady while active annotators diverged.
VI. DISCUSSION & CONCLUSIONS
The paper frames inter-annotator disagreement as structured behaviour shaped by construct abstraction, contrasting consensus-dominant sensory annotation with effort-limited abstract appraisal. It concludes that disagreement can be diagnostically and practically informative, while emphasizing that the evidence is associational and limited in scope.
- Conclusions: The Ordinal Annotation Game treats disagreement as a structured behavioural phenomenon shaped by construct abstraction.Sensory judgments produce a consensus-dominant regime, whereas abstract appraisal tasks produce an effort-limited regime.
- Conclusions: Frequent updates reinforce consensus for concrete sensory judgments but reduce consensus for abstract appraisal tasks.The identical annotation pipeline produces a positive sensory slope and an opposite sign in the abstract task.
- Limitations: The conclusions are associational because the corpora differ in complexity and participant population.The paper also reports a small cohort (N2 = 5) and a low-level saliency proxy, limiting statistical power and generalisability.
- Practical implications: The framework provides a diagnostic for distinguishing structured from random disagreement before label aggregation.It motivates construct-aware annotator weighting and treating active off-consensus labels as potentially informed private appraisal rather than inattentiveness.
- Practical implications: A single majority-vote ground truth may be inappropriate in some settings, where a distributional representation of annotation is more faithful.This follows the framework’s interpretation of disagreement as an informative behavioural signal rather than measurement noise.
- Limitations: The findings are bounded by two corpora, small cohorts, and a low-level proxy, limiting generalizability.The authors recommend manipulating construct abstraction within one corpus while holding other factors constant.