Source-linked AI summary

Whose Judgments Count? Representation Gaps in Crowdsourced Content Moderation Produce Unequal Protection from Perceived Toxicity

Zhaodi Chen, Byungkyu Lee

arXiv:2609.01625v1cs.SIcs.CLcs.CY

TL;DR

The paper asks how demographic differences in moderation judgments affect which users are protected from perceived toxicity. It combines large-scale removal judgments with counterfactual simulations of moderator-pool composition and finds that protection is redistributed toward groups represented in the pool, while representative pools still leave Black and LGB users underprotected.

  • Problem

    It remains unclear how aggregating human moderation judgments shapes protection across users when moderation demand differs demographically.

  • Method

    The study links a multi-identity analysis of moderation demand with counterfactual simulations varying moderator-pool demographic composition.

  • Results

    Aggregated differences in what groups consider removable can produce unequal protection, with representation changing the distribution of protection across groups.

  • Takeaways & Limitations

    More representative judgment pools reduce protection gaps, but descriptive representation alone does not ensure equal protection for Black and LGB users.

  • Takeaways & Limitations

    The findings’ generalizability is constrained because the sample represents U.S. perspectives.

Abstract

from arXiv · show

Content moderation is a central form of digital governance, yet people disagree over what content should be removed from shared online spaces. While platforms aggregate human judgments to build moderation systems, it remains unclear how this process shapes which users are protected from content they perceive as toxic. We address this gap by combining large-scale judgment data with counterfactual simulations that trace how the demographic composition of moderator pools shapes the distribution of protection across users. Applying this framework to removal judgments from 16,221 U.S. respondents evaluating 102,463 comments from Twitter, Reddit, and 4chan, we find demographic heterogeneities in moderation demand. We further reveal a consistent pattern of in-group protection: reductions in perceived toxicity accrue disproportionately to users who share the demographic identities of the moderator pool. Crucially, moderator pools that mirror the demographic composition of self-identified moderators on Prolific widen these disparities relative to a nationally representative baseline, while even fully representative pools fail to ensure equal protection: Black and LGB users remain underprotected unless they are represented well beyond their population share. These findings show that unequal protection from perceived toxicity can arise structurally from the aggregation of stratified removal standards, making the demographic composition of moderation inputs a key determinant of who is protected online.

Results

Moderation demand differs across demographic groups, and changing moderator-pool composition redistributes perceived-toxicity protection. Across simulated and empirically grounded pools, in-group representation generally increases protection for that group, while Black and LGB users remain persistently underprotected.

  • Women show greater moderation demand than men, with predicted probabilities of 0.33 versus 0.30.
  • Asian participants exhibit the highest moderation demand at 0.39, followed by Black participants at 0.35 and Hispanic participants at 0.33.
  • LGB respondents have higher removal demand than non-LGB respondents, at 0.33 versus 0.31, and these differences are not reducible to perceived-toxicity differences.
  • Increasing a group’s moderator-pool representation generally shifts relative toxicity reduction toward that group, producing an in-group protection effect.The focal group’s protection rises while comparison-group protection generally declines or grows more slowly.
  • LGB users remain less protected than non-LGB users until LGB representation reaches roughly 70%, while Black users remain less protected than White users until Black representation reaches roughly 60%.
  • Empirically grounded pools widen protection gaps relative to nationally representative pools, while proportional representation narrows but does not close them.Black and LGB users remain least protected even under representative conditions.

Discussion

The discussion argues that demographic composition is a governance choice with distributional consequences: aggregating stratified moderation standards can produce unequal protection from perceived toxicity. Representation reduces some protection gaps, but Black and LGB users remain under-protected under representative conditions, while the findings are bounded by U.S.-specific data and simplified simulations.

  • Women, racial minorities, and LGB individuals are more likely than counterparts to judge identical comments as warranting removal.
  • These identity-linked moderation demands can be translated through aggregation into unequal protection across groups.
  • Changing whose removal judgments enter moderation changes whose perceived toxicity is reduced, making judgment-pool composition a distributional governance choice.
  • Demographically skewed pools can channel greater protection toward dominant groups, consistent with in-group protection.
  • More representative judgment pools reduce protection gaps, but Black and LGB users remain persistently under-protected under representative conditions.
  • The findings are limited by U.S.-only perspectives, potentially nonrepresentative professional moderators, platform-specific comment coverage, and simulations that omit complex review workflows and contextual factors.

Materials and Methods

The study combines a large U.S. toxicity-judgment dataset with counterfactual simulations that vary moderator demographics and evaluate alternative moderation rules.

  • Scope: The sample is not nationally representative and differs from the 2024 GSS on age, education, and the shares of White and Black respondents.These differences constrain direct generalization from the analytical sample.
  • Data: The analytical sample contains 16,221 adult respondents, 102,463 comments, and 501,540 ratings from Twitter, Reddit, and 4chan.Each comment received about five ratings on average; comments likely to be toxic or contested were oversampled.
  • Data: The analysis excludes respondents with missing demographic responses and nonbinary respondents because their number was too small for composition simulations.The excluded nonbinary respondents represented about 1% of the sample.
  • Simulation design: The framework varies hypothetical 20-person moderator pools while holding other institutional factors constant, repeating each simulation 2,000 times.Raking targets demographic compositions while limiting confounding from correlated characteristics.
  • Modeling and outcomes: Alternative decision rules combine moderator judgments using majorities, a supermajority, or an any-flag rule, with unobserved decisions estimated by an annotator-conditioned model.The model uses comment text, toxicity ratings, rater-specific embeddings, and respondent demographics; post-moderation toxicity weights comments by their probability of remaining visible.
  • Simulation design: Scenarios vary individual demographic attributes or match empirical profiles from the general U.S. population, Reddit moderators, and Prolific respondents with or without moderation experience.The Reddit-matched scenario cannot independently constrain racial or LGB composition because published Reddit moderator data lack those margins.
  • Outcomes: The study measures absolute toxicity reduction and relative toxicity reduction, defined respectively as the mean toxicity difference and that reduction divided by pre-moderation toxicity.Estimates are means across 2,000 runs, with 95% Monte Carlo confidence intervals when shown.
  • Robustness: Robustness analyses recode context-dependent judgments as removal and replace global removal with respondents’ personal filtering decisions.The primary specification counts only “should be removed” as removal, treating conditional approval as insufficient.

Supporting Text

The supporting analyses formalize group-specific toxicity outcomes and model respondent-level removal decisions, while calibrating imputed decisions to observed group relationships.

  • Statistical models: Comment fixed-effects logistic regressions isolate respondent-characteristic differences in moderation demand while holding comment content, tone, and context constant.Controls include race, gender, LGB status, education, age, and political affiliation.
  • Outcome definitions: The simulation defines pre-moderation toxicity as each group’s mean perceived toxicity and post-moderation toxicity using comments that remain after removal.Each comment is assigned a binary moderation decision, with 1 indicating removal.
  • Outcome definitions: Absolute Toxicity Reduction is the difference between mean perceived toxicity before and after moderation, while Relative Toxicity Reduction divides that difference by pre-moderation toxicity.Both measures are computed separately for each demographic group.
  • Decision imputation: The annotator-conditioned model predicts respondent-specific removal probabilities from comment text, toxicity-domain scores, respondent embeddings, and demographics.A single model is trained across respondents because smaller groups lack enough ratings for independent models and shared signal would otherwise be discarded.
  • Validation: The model achieves AUC-ROC = 0.84 for individual removal decisions, versus 0.74 for a text-only model and 0.66 for Perspective scores alone.Low inter-rater agreement was Krippendorff’s α = 0.14; group removal rates were reproduced within approximately one percentage point.
  • Calibration: On the most toxic comments, imputed between-group gaps are roughly half their observed size because unobserved individual toxicity perceptions limit prediction.Across six alternative architectures, unobserved perceived toxicity could be predicted only to approximately 0.6 points on the 0-4 scale.
  • Simulation implementation: The simulation uses imputed probabilities for missing decisions, preserves observed decisions where available, and computes each group’s toxicity outcomes only from observed group ratings.The group-level calibration restores coupling to observed rates, not to specific individuals.

SI Figures

The supplementary figures document the simulation workflow, validate imputed moderation decisions, and examine how moderator-pool composition and decision rules affect relative toxicity reduction.

  • Alternative decision rules: The supplementary simulations reproduce relative toxicity reduction across moderator-pool compositions under a majority-of-five removal rule applied to all 102,463 comments.A comment is removed when at least three of five assigned moderators would remove it, with estimates averaged across 2,000 runs per scenario.
  • Simulation procedure: Simulation weights are constructed by matching target demographic margins before repeatedly sampling moderator teams and recording their decisions.Counterfactual scenarios resample baseline data to impose target group proportions, while empirical-benchmark scenarios use observed survey-weighted margins.
  • Validation: Imputed moderation decisions preserve the observed group-specific relationship between perceived toxicity and removal, recovering 69% to 81% of observed gaps in the highest-toxicity bin.The imputed data cannot fully recover within-comment individual variation because it extends decisions to respondents who did not rate each comment.
  • Empirical benchmark scenarios: Empirically grounded scenarios compare general-population, moderation-experience, no-experience, and Reddit-matched moderator pools using relative toxicity-reduction estimates.The figures distinguish mean pre-moderation toxicity from post-moderation toxicity and annotate relative toxicity reduction across 2,000 runs.
  • Personal filtering preference: Alternative simulations evaluate predicted moderation demand and group-specific relative toxicity reduction when decisions use personal filtering preferences rather than the standard removal rule.The plotted outcomes are reported by gender, race, and LGB status with 95% confidence intervals or error bands.

SI Tables

The supplementary tables define the analytic sample and decision rules, report regression and protection-gap results, and document demographic coverage and simulation diagnostics.

  • Sample and measures: The analytic sample contains 501,540 respondent–comment ratings from 16,221 respondents evaluating 102,463 comments.Binary moderation recodes “This comment should be removed” as moderated and the other listed responses as not moderated.
  • Regression models: Regression estimates model moderation demand and toxicity perception with comment fixed effects, using 236,645 ratings from 16,172 respondents across 47,553 comments for moderation-demand models.Comments without variation in binary moderation judgments do not identify conditional fixed-effects logistic regression.
  • Protection gaps: Increasing a focal group’s moderator-pool share from 10% to 90% changes its toxicity-reduction gap relative to a comparison group, with positive changes indicating movement toward the focal group.Reported values are percentage points, and gaps are defined as focal group minus comparison group.
  • Benchmark pools: Reddit-matched simulations rake only to gender, age, and education margins because published Reddit moderator data do not report race or LGB identity.This creates a narrower demographic match than scenarios using the fuller set of available margins.
  • Coverage and diagnostics: Observed rating coverage is uneven across groups, with women, men, White respondents, and non-LGB respondents rating nearly all comments while some groups have substantially lower coverage.Each rating record contains both toxicity and moderation information for the same respondent–comment pair.
  • Coverage and diagnostics: Raking and sampling generally reproduce intended moderator-pool compositions closely, although sparse intersecting demographic cells produce the largest deviations in Asian-share scenarios.Raking weights were capped at 10 to avoid assigning extreme influence to sparse demographic cells.
Loading 2609.01625v1…