Source-linked AI summary

Community corrections have divergent downstream effects across corrected accounts

Yuwei Chuai, Thomas Renault, Nicolas Pröllochs, Gabriele Lenzini, Mohsen Mosleh

arXiv:2608.27526v1cs.SI

TL;DR

The study asks whether community-based fact-checking changes corrected authors’ later behavior, beyond reducing the spread of individual annotated posts. Using a large-scale quasi-experimental Difference-in-Differences design on X, it finds a 2.9% average increase in original-post activity, masking divergent responses between accounts corrected once and repeatedly. These results distinguish correcting content from changing the behavior of its producers.

  • Problem

    Whether community notes induce behavioral changes extending beyond the corrected post remains largely unknown, despite evidence of post-level effects.

  • Method

    The study shifts analysis from corrected posts to authors’ subsequent behavior using a quasi-experimental design tracking accounts before and after corrections.

  • Results

    2.9% more posts were published during the four weeks after correction on average, with accounts corrected once reducing activity and repeatedly corrected accounts increasing it.

  • Takeaways & Limitations

    Correcting misleading content and improving producer behavior are distinct objectives for platform governance.

  • Takeaways & Limitations

    The quasi-experimental design cannot exclude unobserved time-varying confounding.

Abstract

from arXiv · show

Community-based fact-checking can reduce the spread of annotated misleading posts, but whether it produces lasting behavioral change among corrected authors remains unclear. Here, we conduct a large-scale quasi-experimental study of the Community Notes system on X (formerly Twitter), tracking four weeks of activity before and after note display for 19,854 accounts and 57,935 corrections (noted posts and matched controls), covering 11,909,591 original posts. Difference-in-Differences estimates show that note display is followed by an average 2.9% increase in corrected accounts' original-post activity. This aggregate conceals two divergent trajectories. Accounts corrected only once reduce their activity by 2.4% and subsequently publish less toxic and less misleading content. Repeatedly corrected accounts, which constitute 28.6% of corrected accounts but produce 73.4% of fact-checked posts, instead increase their activity by 4.4% after their first correction and show no detectable response to later ones. They exhibit no comparable content improvement, and instead publish more highly misleading posts, cite lower-quality domains, and post more political content. Community notes can thus constrain individual misleading posts without durably improving the behavior of the accounts most responsible for them, indicating that correcting content and changing its producers are distinct objectives for platform design.

Introduction

Community-based fact-checking can reduce the impact of misleading posts, but its effects on the authors who produce them remain unclear. This study examines whether Community Notes change corrected authors’ subsequent activity and content, distinguishing accounts corrected once from those corrected repeatedly.

  • Prior evidence: Community-based fact-checking can reduce reposting, engagement, diffusion, and persistence of annotated misleading posts.Prior quasi-experimental and experimental research reports reduced engagement, altered diffusion, increased deletion, and improved identification of misleading content after notes or contextual corrections.
  • Research gap: Whether these post-level effects extend to authors’ later posting behavior and content quality remains largely unknown.Deleting or interrupting a corrected post does not necessarily mean its author posts less or produces more reliable content.
  • Why author behavior matters: Author-level responses matter because suppressing individual posts and changing producers’ behavior are distinct objectives.This distinction is especially consequential when a small minority of accounts disproportionately contributes to misleading information circulation.
  • Competing expectations: Community Notes could encourage caution through corrective information and reputational costs, but public corrections might also provoke resistance.The introduction presents competing expectations: authors may reassess posting practices, or respond defensively and shift toward social or partisan considerations.
  • Study design: The study uses a large-scale quasi-experimental Difference-in-Differences design to compare four weeks before and after note display across correction-frequency groups.It tracks 19,854 corrected accounts and 11,909,591 original posts, measuring both posting activity and content while distinguishing accounts corrected once from repeatedly corrected accounts.
  • Preview of findings: Accounts corrected once reduce activity and publish less toxic, misleading, and political content, whereas repeatedly corrected accounts increase activity without improving content quality.The study concludes that community notes do not improve, on average, the subsequent behavior of the accounts most responsible for fact-checked misleading posts.

Results

Note display is followed by a sustained increase in corrected accounts’ posting activity, but responses diverge by correction frequency. Accounts corrected once reduce activity and improve content, whereas repeatedly corrected accounts increase activity after their first correction without comparable improvement.

  • Overall activity response: 2.9%: Corrected accounts increased posting activity relative to matched controls across the four post-display weeks.The increase remained positive in each post-display week and persisted throughout the four-week period.
  • Heterogeneity by correction frequency: 2.4%: Accounts receiving only one note reduced their posting activity after correction.The estimate was −0.024 relative to matched controls.
  • Heterogeneity by correction frequency: 4.4%: Accounts subsequently receiving multiple notes increased activity after their first correction, with larger increases among accounts receiving at least three, four, or five notes.The corresponding increases were 6.9%, 8.1%, and 9.5%, respectively.
  • Timing of repeated-account responses: Later corrections produced no statistically detectable additional activity changes among repeatedly corrected accounts.Estimates from the second correction onward were statistically indistinguishable from zero.
  • Pre-correction differences: Repeatedly corrected accounts differed from single-correction accounts before correction in posting frequency, sentiment, misleadingness, political content, and account characteristics.They posted more frequently and had more negative sentiment, misleading content, opinion confidence, and political content, alongside less positive sentiment and lower toxicity.
  • Content response: Accounts corrected once subsequently published less toxic and less misleading content, whereas repeatedly corrected accounts showed no comparable reductions.Single-correction accounts reduced high-toxicity and low-URL-quality posts; repeatedly corrected accounts published more highly misleading posts, cited lower-quality domains, and posted more political content.

Discussion

Community Notes produce divergent author-level responses: one-time-corrected accounts reduce activity and improve content, while repeatedly corrected accounts increase activity without comparable improvement. These findings distinguish constraining individual misinformation from changing the behavior of its producers.

  • Accounts corrected only once reduced posting activity and subsequently published less toxic and less misleading content.This pattern supports different behavioral trajectories after correction frequency is taken into account.
  • 2.9% more original posts followed note display on average, but this correction-event estimate was disproportionately shaped by repeatedly corrected accounts.Such accounts were 28.6% of corrected accounts yet contributed 73.4% of fact-checked posts.
  • 4.4% higher activity followed the first correction among accounts later corrected repeatedly, while later corrections produced no statistically detectable response.The divergent response was already apparent after the first observed correction rather than emerging after multiple corrections.
  • Repeatedly corrected accounts already posted more frequently and produced more negative, misleading, and political content before correction than one-time-corrected accounts.They also had more followers and were more likely to be verified, right-leaning, and exposed to misinformation; these between-group differences do not establish causation.
  • Repeatedly corrected accounts showed no detectable reduction in toxicity or misleadingness, cited lower-quality domains, posted more political content, and produced more highly misleading posts.Their average misleadingness did not significantly increase, but greater activity increased the total volume of highly misleading posts.
  • The first correction may be an important intervention point, whereas repeatedly displaying the same correction may be insufficient for frequently corrected accounts.The authors suggest evaluating author-directed interventions alongside content correction while treating content correction and behavioral change as distinct governance objectives.

Methods

The study uses post-level treatment and matched control events to examine how Community Notes affect corrected authors’ subsequent activity and content. It combines longitudinal activity counts with post-level content measures in Difference-in-Differences analyses over four-week pre- and post-display windows.

  • Data and design: Four-week windows before and after each actual or assigned note-display time define the observation periods.The activity analysis covers 28 days before and 28 days after display for each post.
  • Content measures: 11,909,591 original posts are retrieved for content analyses, while the main analysis focuses on authors’ subsequent original posts.Community Notes predominantly fact-check original posts, with only 3.8% of fact-checked posts being replies.
  • Statistical analysis: The analysis accounts for repeated corrections and excludes overlapping treatment and control windows to avoid simultaneous attribution of account activity.Treatment assignment remains specific to individual posts and their observation windows.
  • Statistical analysis: Difference-in-Differences models compare posting changes after note display with matched controls while estimating temporal leads and lags.The leads assess parallel trends, and the post-treatment lag coefficients measure weekly or daily impacts.

Data and code availability

The authors state that pseudonymized data, materials, and analysis scripts required to reproduce the study will be made publicly available upon publication.

  • Pseudonymized data, materials, and analysis scripts required for reproduction will be made publicly available upon publication.

Supplementary Materials

The listed authors are Yuwei Chuai, Thomas Renault, Nicolas Pröllochs, Gabriele Lenzini, and Mohsen Mosleh.

  • The paper lists five authors: Yuwei Chuai, Thomas Renault, Nicolas Pröllochs, Gabriele Lenzini, and Mohsen Mosleh.

S1 Data overview

The dataset covers corrected and matched control posts from thousands of accounts, with most corrected accounts receiving one note but repeatedly corrected accounts producing most fact-checked posts.

  • 57,935 community note corrections cover 19,854 accounts, including 29,049 treated posts from 10,799 accounts and 28,886 control posts from 12,823 accounts.
  • Most accounts whose misleading posts were corrected received only one community note.
  • 11.1% of treated accounts received at least two community notes, while 8.6% received at least three.
  • 21,336 fact-checked posts, or 73.4%, originated from accounts receiving at least two community notes.
  • The data overview reports counts, means with standard deviations, and percentages for account, posting, sentiment, toxicity, URL-quality, and political-content measures.

S2 Robustness checks for changes in post count

Robustness analyses consistently support a positive post-correction change in corrected accounts’ original-post activity, despite alternative windows, fixed effects, treatment definitions, and sensitivity assumptions.

  • 2.9% is the aggregated weekly increase in original-post activity after corrections, consistent with the 2.7% aggregated daily increase.
  • Pre-correction daily and weekly estimates are not statistically distinguishable from zero, whereas several post-correction estimates are significantly positive.
  • Post-correction weekly effects rise from 1.1% in Week 1 to 2.8%, 3.0%, and 3.7% across Weeks 2–4.
  • Post-level fixed-effects estimates remain positive at 2.6% over four weeks and are not significantly different from the main random-effects estimates.
  • HonestDiD sensitivity analysis preserves the positive effect under deviations up to 1.0 times short-term and 1.5 times long-term pre-trend violations.
  • The alternative estimator accommodates repeated, varying correction exposure and produces consistent results under binary and non-binary treatment specifications.

S3 Heterogeneity by correction frequency

Table S4 reports group sizes for the estimations presented in Figure 2, whose panels distinguish correction-frequency analyses.

  • Figure 2: Table S4 reports group sizes for the estimations in Figure 2.The table is identified as supporting the Figure 2 estimations.
  • Figure 2: Figure 2 includes panels 2b, 2c, and 2d for the correction-frequency analyses.The supplied panel labels identify these three Figure 2 panels.
  • Table S4: The table headings distinguish corrections, observations, and events across the reported analyses.The headings are repeated across the displayed columns.

S4 Heterogeneity by account profiles and posting preferences

Responses to first correction vary across pre-correction profiles and posting preferences, with the largest increases among accounts already posting more toxic, misleading, or political content.

  • Between-group characteristics associated with repeated correction need not predict stronger within-group responses to a first correction.
  • First-correction activity increases across account-characteristic subgroups, rather than being confined to a particular profile level.
  • 12.3% is the activity increase for low-frequency accounts after their first correction, while high-frequency accounts show a nonsignificant 1.2% estimate.
  • 5.4% is the significant activity increase among accounts with high prior toxicity, compared with a nonsignificant 2.2% estimate for low-toxicity accounts.
  • 5.8% is the activity increase for accounts with high prior misleadingness, while the low-misleadingness estimate is a nonsignificant 0.8%.
  • 5.8% is the activity increase among accounts with high prior political content, compared with a nonsignificant 0.8% estimate for less-politically oriented accounts.

S5 Parallel-trends assessment for post content outcomes

Figure S8 assesses parallel trends for post-content outcomes using leads-and-lags and aggregated Difference-in-Differences estimates. Pre-display estimates are statistically indistinguishable from zero, supporting the parallel-trends assumption.

  • Reference periods: Week −1 is the reference period for most estimates, while daily misleadingness uses Day −1.
  • Parallel-trends assessment: Pre-display estimates are statistically indistinguishable from zero, supporting the parallel-trends assumption across outcomes.
  • Estimation groups: Tables S6 and S7 report group sizes for the estimations shown in Figures 4 and 5, respectively.
  • Content-outcome estimates: Figure S8 reports weekly, daily, and aggregated estimates for sentiment, toxicity, URL quality, misleadingness, confidence, and politics outcomes.The figure includes positive and negative sentiment, toxicity, URL quality, misleadingness, confidence, and politics.
Loading 2608.27526v1…