Source-linked AI summary

Beyond Effectiveness: A Multi-Criteria Framework for Comparing Practical Socio-Technical Interventions

Catherine King, Lynnette Hui Xian Ng, Kathleen M. Carley

arXiv:2608.20649v1cs.AI

TL;DR

Sociotechnical interventions are often evaluated individually and mainly for effectiveness, leaving practitioners without a principled basis for comparing practical trade-offs. This paper develops a five-criteria framework and finds that experts’ most effective interventions are not always the most acceptable or feasible.

  • Problem

    Sociotechnical interventions are often evaluated individually and mainly on effectiveness, rather than comparatively across practical and normative criteria.

  • Method

    The PEACE framework compares 40 misinformation interventions across political feasibility, effectiveness, acceptance, cost, and implementation effort using a survey of 39 experts.

  • Results

    Experts’ most effective interventions are often not those judged most acceptable or feasible, with restrictive interventions facing barriers and empowering interventions showing uncertain effectiveness.

  • Takeaways & Limitations

    The framework supports principled intervention decisions by identifying trade-offs among effectiveness, acceptance, feasibility, cost, and effort.

  • Takeaways & Limitations

    The expert sample was small, geographically limited, and predominantly drawn from a Northern American academic institution.

Abstract

from arXiv · show

Designers and policymakers in sociotechnical domains like content moderation, privacy interfaces, recommender systems and beyond, must choose among a growing menu of proposed interventions, but typically lack a principled basis for comparing them. Prior work tends to evaluate interventions individually and mostly along the effectiveness criteria, while implementation constraints such as cost, effort and feasibility are often considered separately. We present a multi-criteria framework for evaluating sociotechnical interventions. This framework is instantiated through the case of misinformation, a domain of intense focus for proposed countermeasures. We survey $N=39$ researchers on 40 operationalized interventions across five evaluative criteria: political feasibility, effectiveness, user acceptance, cost, and implementation effort. We find that the interventions that experts judge to be the most effective are not always the most acceptable to the public or the most feasible to implement. We also discuss how this tension has implications for the design of sociotechnical interventions beyond misinformation, and offer a decision framework for practitioners navigating the trade-offs of sociotechnical interventions.

1 Introduction

The paper develops a multi-criteria framework for comparing socio-technical interventions, using misinformation to examine trade-offs among political feasibility, effectiveness, acceptance, cost, and effort. It introduces a taxonomy and expert evaluation to guide practical intervention design.

  • Motivation: Socio-technical interventions are often developed and evaluated individually rather than compared systematically across standardized criteria and social contexts.The paper argues that interventions cannot be meaningfully evaluated outside the social context in which they are deployed.
  • Framework: The study develops a multi-criteria framework through misinformation, where countermeasures are proliferating faster than the field’s capacity to compare them.The paper defines an intervention as a deliberate policy, platform design, or educational effort deployed to reduce a socio-technical harm.
  • Research questions: The framework investigates trade-offs across political feasibility, effectiveness, user acceptance, cost, and implementation effort.Its research questions also ask which intervention categories balance user-centered and implementation-centered criteria and what design principles can guide practical design.
  • Contributions: The paper introduces a taxonomy of forty operationalized interventions organized into eight categories, including moderation, distribution, labeling, user-based measures, education, and institutional measures.The categories also include generative AI-specific measures and other external and institutional measures.
  • Contributions: It provides expert evaluation of the interventions across five criteria: political feasibility, effectiveness, user acceptance, cost, and implementation effort.The expert elicitation survey examines how interventions vary across these criteria.

2 Related Work

Prior work on misinformation interventions typically emphasizes effectiveness or user acceptance and rarely compares these dimensions alongside implementation constraints. This framework addresses that gap by evaluating interventions across multiple criteria and accounting for both objective features and user perceptions.

  • Prior Reviews: Prior reviews categorize misinformation interventions by platform component or implementing actor, but primarily evaluate efficacy while only suggesting feasibility, scalability, and durability as additional dimensions.
  • Prior Reviews: Existing reviews rarely evaluate effectiveness and user acceptance together, and almost never alongside implementation constraints.
  • Framework Gap: The framework compares multiple intervention dimensions at once to help identify appropriate interventions for a given context.
  • Intervention Features: Interventions comprise objective implementation characteristics and potentially changing user perceptions that influence policy support.
  • Cross-Domain Foundations: HCI and public policy research frame sociotechnical interventions as trade-offs among competing criteria rather than single-target optimization.

3 The PEACE Framework

The PEACE framework evaluates sociotechnical interventions across political feasibility, effectiveness, user acceptance, cost, and implementation effort. It organizes misinformation interventions around the social-media information lifecycle, spanning creation, dissemination, and correction or prevention.

  • Evaluative criteria: PEACE assesses interventions using five criteria: political feasibility, effectiveness, acceptance, cost, and effort.The criteria were selected from prior comprehensive reviews and policy analyses.
  • Intervention categorization: The information-lifecycle lens distinguishes creation, dissemination, and engagement through correcting, contesting, or contextualizing information.This lens identifies where analysts could intervene on social-media information flows.
  • Intervention categorization: Six intervention categories operate directly on platforms, while media literacy and external institutional measures occur primarily beyond platforms.Platform interventions include account moderation, content moderation, content distribution, generative AI-specific measures, content labeling, and user-based measures.
  • Evaluative criteria: Political feasibility incorporates stakeholder perspectives, regulatory and legal constraints, and support from relevant officials or organizations.Effectiveness concerns success in targeting creation, spread, or belief phases under different circumstances.
  • Decision framework: The framework treats intervention viability as a balance shaped by inherent characteristics and user perceptions across the five criteria.User perceptions can affect acceptance, public policy support, compliance, and effectiveness.

4 Data and Methods

The study used an expert elicitation survey to compare misinformation interventions across five evaluative criteria. Thirty-nine researchers assessed intervention categories and randomly assigned operationalized interventions using Likert scales, with cost and effort scores inverted for consistent interpretation.

  • Participants and survey design: Thirty-nine researchers completed the survey between January 23 and March 24, 2025.Participants were recruited through SUMMIT 2025 participants and invitees.
  • Participants and survey design: Experts rated general intervention categories on effectiveness and acceptability using a 1–7 Likert scale, then evaluated 12 of 40 operationalized interventions across five metrics using a 1–5 scale.The five metrics were political feasibility, effectiveness, user acceptance, cost, and effort level; each participant’s 12 interventions were randomly selected.
  • Measurement and analysis: Cost and effort scores were inverted so higher values indicated better results across all five metrics, and the measures were renamed cheapness and ease of implementation.A score of 1 indicated low levels and 5 high levels before inversion; after inversion, 5 indicated low cost and effort.

5 Results and Analysis

Experts’ evaluations reveal persistent trade-offs among effectiveness, acceptance, feasibility, cost, and implementation effort. Clustering further distinguishes interventions that are balanced, effective but contested, effective but costly, or easy to implement but low-impact.

  • Category-level results: Content moderation, account moderation, and media literacy / education were regarded as the most effective categories, while account moderation and user-based measures showed the largest effectiveness–acceptance gaps.The ratings compare experts’ judgments of effectiveness with acceptance among the American public.
  • Multi-criteria evaluation: Experts rated individual interventions across political feasibility, effectiveness, user acceptance, cheapness, and ease of implementation, with content labeling, user-based measures, and content distribution leading category averages.Cheapness and ease of implementation are inverted scores for cost and effort.
  • Metric trade-offs: Only three of the top ten overall interventions also ranked in the top ten for effectiveness, while five did so for acceptance, effort, and cost, and seven for political feasibility.Most top interventions ranked in the top ten for only two or three metrics.
  • Intervention clusters: The four clusters were Balanced High Performers, Effective but Contested, Effective but Costly, and Easy Starters, containing 12, 15, 9, and 4 interventions respectively.These groups organize interventions by similar criterion scores and expose distinct trade-offs across the five metrics.
  • Intervention clusters: Balanced High Performers scored relatively well across all five metrics, whereas Effective but Contested interventions were low-cost and effective but had relatively low user acceptance.The balanced cluster included less intrusive or higher-agency measures, while the contested cluster generally included restrictive platform controls.
  • Intervention clusters: Effective but Costly interventions were acceptable and effective but resource-intensive, while Easy Starters were politically feasible and popular but generally viewed as less effective.The costly cluster included institutional and educational measures; Easy Starters had the highest overall user acceptance and implementation likelihood.

6 Discussion

The discussion emphasizes systematic trade-offs among feasibility, effectiveness, acceptance, cost, and effort, with user agency and participatory input shaping more sustainable interventions. It extends these implications beyond misinformation by highlighting lifecycle coverage and the need to address creation-phase incentives.

  • Discussion: Interventions involve systematic trade-offs across political feasibility, effectiveness, acceptance, cost, and effort rather than a single optimal solution.This pattern recurs across both category-level and individual-level interventions.
  • Design implications: Informational interventions that preserve user agency generally outperform more intrusive or restrictive categories across the five evaluation metrics.Content labeling, user-based measures, and content distribution scored highest on average among expert raters.
  • Design implications: Expert judgments should not substitute for public input, because experts and the public can disagree about the intrusiveness and acceptability of interventions.The discussion recommends participatory design to surface disagreements before deployment and make online spaces more inclusive.
  • Beyond misinformation: Across sociotechnical systems, practitioners should combine interventions across mechanisms and design creation-phase measures that restructure incentives without restricting participation.Account and content visibility, distribution, and moderation mechanisms recur beyond misinformation, including in coordinated inauthentic behavior.
  • Design implications: Acceptance should remain a design target for highly effective interventions, with fairness, transparency, trust, recourse, and source transparency helping reduce opposition.Account and content moderation were rated effective by 74% and 77% of expert participants, respectively, but were less acceptable.
  • Information lifecycle: The most viable interventions cluster around misinformation’s spread and belief phases, leaving creation-phase incentives comparatively underaddressed.Only two of the top 10 interventions directly target creation incentives: altering platform metrics and demonetization; both average 3.75/5 for acceptance and rank 19th/20th.

7 Conclusion

The paper presents a multi-criteria framework for comparing sociotechnical interventions, instantiated through a survey of misinformation interventions. Its findings show that effectiveness can conflict with acceptability, cost, and political feasibility, motivating principled trade-off decisions.

  • Framework and study design: N=39 researchers evaluated 40 misinformation interventions across five criteria within eight intervention categories spanning the information lifecycle.The criteria included political feasibility, effectiveness, user acceptance, cost, and implementation effort.
  • Empirical findings: The interventions judged most effective were often not those judged most acceptable or politically feasible.This result demonstrates why evaluating interventions on effectiveness alone is insufficient.
  • Empirical findings: User-based measures, content labeling, and content distribution measures scored highly across multiple evaluative metrics.These intervention types performed strongly on more than one criterion in the expert evaluations.
  • Empirical findings: Provably effective content and account moderation techniques can lack user support because they are perceived as unfair or intrusive.Perceived unfairness and intrusiveness may undermine acceptance despite demonstrated effectiveness.
  • Implications: The framework is intended to support principled design decisions by identifying trade-offs that can otherwise cause highly effective interventions to fail in deployment.Deployment can fail when an intervention is unacceptable, too costly, or politically infeasible.
  • Implications: The proposed comparative framework is a first step that can extend beyond misinformation to other sociotechnical intervention domains.The paper positions the framework as applicable to broader design and governance decisions.

A Intervention Categorization

Four review articles classify misinformation interventions by different dimensions, including the targeted misinformation driver, information-lifecycle phase, platform component, and intervention function. Because existing categories overlap and lack a common typology, the paper synthesizes them into eight general countermeasure categories.

  • Literature basis: Four prominent reviews provide the literature basis for the intervention categorization taxonomy.They span HCI, psychology, interdisciplinary social science, and Nature Human Behaviour venues.
  • Existing taxonomies: The reviews classify interventions by platform component, misinformation driver, lifecycle phase, or behavioral function.Dimensions include content, sources, individual users, communities, creation, spread, belief, prevention, nudges, boosts, education, and refutation.
  • Existing taxonomies: The literature includes categories such as content labeling, moderation, reporting, literacy, redirection, security, fact-checking, warning labels, and source-credibility labels.Fact-checking is included under disinformation disclosure in one review, while redirection and credibility labels are treated as forms of content labeling elsewhere.
  • Taxonomy gap: Review categories overlap, and no common typology exists across the literature.Examples include lateral-reading skills versus media literacy, while user-led and institutional interventions such as government regulation are often absent.
  • Proposed categorization: The paper synthesizes the four reviews into eight general countermeasure categories.These categories are primarily organized by the targeted information-lifecycle phase and affected platform component, with detailed intervention subcategories described subsequently.

A.1 Account Moderation

Account moderation consists of platform interventions targeting user accounts, including suspension, removal, visibility limits, and demonetization. These measures can temporarily or permanently restrict accounts or suppress their reach while leaving them active.

  • A.1 Account Moderation: Account moderation targets user accounts through suspension, removal, shadowbanning, or demonetization, restricting access, visibility, or monetization.Shadowbanning limits an account’s visibility to others, while demonetization restricts its ability to monetize content.
  • A.1 Account Moderation: Account suspension temporarily suspends or permanently bans accounts that repeatedly violate platform policies, while deplatforming removes high-risk accounts across multiple platforms.Deplatforming is coordinated across platforms rather than limited to a single platform.

A.2 Content Moderation

Content moderation encompasses interventions that alter how social-media content is evaluated, presented, distributed, or displayed. These interventions may involve human moderators, users, or automated systems, including detection, fact-checking, and debunking.

  • Scope: Content moderation covers interventions that evaluate information accuracy or modify content visibility and circulation on social media.These interventions concern how content is presented, distributed, or displayed.
  • Intervention types: Misinformation detection uses computational methods to identify potentially false, misleading, or unreliable content for subsequent moderation.Detected content may later be reviewed, labeled, downranked, or removed.
  • Intervention types: Fact-checking assesses claims against available evidence, whereas debunking adds evidence, explanations, or context to correct misleading claims.Fact-checking can be conducted by experts, journalists, and platforms in text or video formats.

A.3 Content Distribution … A.8 Institutional Measures

The paper organizes misinformation interventions into six categories spanning content distribution, generative AI, labeling, user empowerment, education, and institutional action. These categories alter information pathways, add contextual guidance, build user capabilities, or coordinate broader institutional responses.

  • A.3 Content Distribution: Content distribution interventions shape how information circulates, reaches audiences, and is encountered without directly removing or modifying content.Examples include forwarding limits, resharing constraints, sharing delays, and redirects to vetted information.
  • A.4 Generative AI-specific: Generative AI-specific interventions use AI-generated rebuttals, educational initiatives, or conversational chatbots to counter misinformation and false beliefs.AI-related prohibitions, advertising restrictions, and disclosure requirements remain classified under existing moderation, distribution, or labeling categories.
  • A.5 Content Labeling: Content labeling attaches disclosures, warnings, source information, credibility signals, or supplementary context to help users evaluate social-media content.Crowdsourcing shifts labeling from professionals to users whose judgments are aggregated into notes or labels.
  • A.6 User-based Measures: User-based measures empower individuals and communities to shape their information environments through social norms, peer corrections, reporting, blocking, and self-moderation tools.These interventions can encourage responsible sharing, flag harmful posts, reduce unwanted exposure, or support behavioral change.
  • A.7 Media Literacy and Education: Media literacy and education interventions strengthen people’s ability to critically evaluate information and navigate digital sources.They include practical misinformation-identification guidance, literacy assessment, communication strategies, and inoculation or pre-bunking.
  • A.8 Institutional Measures: Institutional measures are undertaken by governments, civil society, news media, and other institutions to govern platforms, support media, and strengthen coordination.Government regulation may include platform accountability laws, privacy rules, antitrust action, or restrictions on micro-targeted advertising.
  • A.8 Institutional Measures: Other institutional measures promote transparency and collective capacity through platform data sharing, civil-society resources, and coordination among relevant institutions.Data sharing supports researchers’ independent investigation through access to raw data and internal findings.

B Operationalized Interventions

This section defines the 40 operationalized misinformation interventions used in the expert survey. Each intervention specifies an implementation drawn from prior literature and organized across eight general categories.

  • Operationalized interventions: The survey covered 40 operationalized misinformation interventions, each representing a specific implementation of one or more interventions from the literature and previous appendix.Participants received definitions of these specific implementations before responding.
  • Operationalized interventions: The interventions were organized across eight general categories.Table 7 presents the operationalized interventions across these categories.
Loading 2608.20649v1…