Source-linked AI summary

SoCRATES: Towards Reliable Automated Evaluation of Proactive LLM Mediation across Domains and Socio-cognitive Variations

Taewon Yun, Hyeonseong Park, Jeonghwan Choi, Hayoon Park, Yeeun Choi, Hwanjun Song

arXiv:2606.05563v1cs.AIcs.CL

TL;DR

LLM mediation is difficult to evaluate because real conflicts vary across domains, emotions, intentions, and context, while existing testbeds provide limited coverage and coarse turn-level scoring. SoCRATES builds realistic scenarios from real disputes, probes five socio-cognitive axes, and evaluates topic-relevant trajectory changes. The evaluator reaches r = 0.82 alignment with experts, while the strongest mediator closes only roughly a third of the unmediated consensus gap; performance varies across socio-cognitive conditions.

  • Problem

    Existing mediation evaluations provide limited domain coverage, vary few socio-cognitive factors, and score trajectories without localizing turns to the topics they advance.

  • Method

    SoCRATES combines agentic curation of real-conflict scenarios, independent probing across five socio-cognitive axes, and topic-localized evaluation.

  • Results

    The strongest mediator closes only roughly a third of the unmediated consensus gap under diverse and realistic testbeds, with performance varying sharply by socio-cognitive axis.

  • Takeaways & Limitations

    Effective LLM mediation depends on adapting intervention strategy to diverse domains, contexts, and party compositions rather than applying uniform behavior.

  • Takeaways & Limitations

    SoCRATES runs conversations in English and focuses primarily on consensus, leaving multilingual mediation and satisfaction, fairness, trust, and emotional repair for future work.

Abstract

from arXiv · show

Evaluating LLM mediators remains challenging, as mediation unfolds as a real-time trajectory shaped by disputants' shifting emotions, intentions, and context. Existing testbeds rely on a few expert-authored domains, vary mainly strategic posture, and score every turn against every topic, introducing off-topic noise. We introduce SoCRATES, a benchmark for evaluating proactive LLM mediators in realistic, multi-domain testbeds. It constructs scenarios from real conflicts through an agentic pipeline across eight domains, probes five socio-cognitive adaptation axes (strategic posture, party composition, history length, emotional reactivity, and cultural identity), and scores each topic only on the turns that advance it via a topic-localized evaluator. The evaluator reaches 0.82 alignment with human experts, more than doubling a per-turn baseline. Benchmarking eight frontier LLMs, we find that even the strongest mediator closes only about a third of the unmediated consensus gap under diverse and realistic testbeds, with performance varying sharply by socio-cognitive axis, highlighting that progress lies in social adaptation to diverse conditions.

1 Introduction

SoCRATES addresses evaluation gaps in proactive LLM mediation by scaling realistic conflict scenarios, independently probing socio-cognitive variation, and localizing trajectory scoring to relevant topics. Its benchmark combines agentic curation, five-axis probing, and topic-localized evaluation to assess mediator adaptation.

  • Existing mediation testbeds cover few expert-authored domains and usually vary only strategic posture, despite conflicts differing in emotion, culture, and history.
  • Mediation evaluation requires scalable scenario coverage, independent socio-cognitive variation, and reliable end-to-end trajectory scoring.
  • SoCRATES realizes this pipeline through agentic scenario curation, socio-cognitive probing, and topic-localized evaluation.
  • Its scenario curation searches public disputes across eight conflict domains, restructures cases, and retains hard scenarios that fail to resolve without mediation.
  • The topic-localized evaluator scores each topic only when turns actively move it and supports consensus gain, intervention timeliness, and intervention effectiveness.
  • The strongest mediator closes roughly a third of the unmediated consensus gap, with gains varying sharply by socio-cognitive axis.

2 Related Work

Related work studies conflict negotiation, third-party mediation, and automated dialogue evaluation, but remains limited by scalability and coarse outcome signals. These limitations motivate scalable, finer-grained evaluation of mediator trajectories.

  • Prior studies use LLMs as disputing parties to reproduce conflict behavior, but this does not reveal how disputes between humans are resolved.
  • Third-party mediation studies require thousands of human disputants to simulate conflicts, creating a scalability bottleneck.
  • Automatic dialogue evaluators often judge negotiation progress through end-state consensus or goal achievement, providing only a coarse view of dialogue state.

3 SoCRATES Framework

SoCRATES represents conflicts as structured multi-party, multi-topic scenarios, generates hard cases from real disputes, independently expands them across socio-cognitive conditions, and evaluates matched trajectories with localized metrics. The framework measures both overall consensus contribution and the timing and effectiveness of interventions.

  • Task Formulation: A conflict scenario is represented as s = (B, P, T, W), combining background, disputants, contested topics, and topic-importance weights.
  • Task Formulation: The mediator observes only shared background, topics, and dialogue, so it must infer hidden party personas, stances, and preferences before deciding when and how to intervene.
  • Agentic Scenario Curation: Agentic curation searches eight conflict domains, recasts web-derived disputes into structured scenarios, and filters candidates through repeated unmediated simulations.
  • Socio-Cognitive Probing: SoCRATES independently applies five socio-cognitive axes to fresh scenario copies, avoiding entangled failures across competencies.
  • Socio-Cognitive Probing: Together with the general condition, the five axes yield 15 conditions covering strategic posture, party composition, history length, emotional reactivity, and cultural identity.
  • Topic-Localized Evaluation: Topic-localized scoring tracks cumulative topic agreement and measures consensus gain, intervention timeliness, and intervention effectiveness.
  • Validation: The evaluator aligns with expert ratings at Pearson r = 0.82, more than doubling both comparison baselines.

4 Validation of SoCRATES

SoCRATES validates both persona controllability and topic-localized evaluation before benchmarking mediators. The evaluator aligns strongly with expert judgments and outperforms per-turn baselines.

  • Simulation Fidelity: Four scalar levels test whether persona intensity produces ordered differences in simulated emotional reactivity.The validation samples scalar values {0, 0.33, 0.66, 1} to assess controllability beyond a binary distinction.
  • Simulation Fidelity: 160 A/B pairs per simulator receive annotations from three crowdworkers, with Krippendorff’s 𝛼= 0.75.DeepSeek-V3.2 achieves the highest score, indicating reliable translation of float-valued persona intensity into ordered reactiveness.
  • Topic-localized Evaluation: Two expert annotators rate 1,844 snippets from 144 mediator trajectories to evaluate trajectory scoring.Snippet-level aggregation preserves evaluator trajectory errors while matching the resolution at which experts can rate reliably.
  • Topic-localized Evaluation: 𝑟= 0.82 on trajectories and 𝑟= 0.80 on outcomes, with the topic-localized evaluator more than doubling both per-turn baselines on trajectories.The comparison uses ProMediate’s every-topic, every-turn judge and a non-expert rater on the same 1–5 scale.

5 Benchmarking LLM Mediators

Across eight domains, LLM mediators achieve limited consensus gains and show uneven adaptation across socio-cognitive conditions. Performance depends on domain, model behavior, and matching intervention timing to the conflict’s demands.

  • Performance by Conflict Domain: Average consensus gain caps at 34.4, and no domain mean clears half the unmediated gap.Mediator means split into top and bottom tiers, while domain means range from 41.3 to 16.6.
  • Performance by Conflict Domain: Proprietary mediators lead the strongest open-source model by 1.1–2.5 points and lead in six of eight domains.Scale helps within a model family, but cross-family rankings do not follow parameter scale consistently.
  • Performance by Conflict Domain: Intervention timeliness alone is insufficient: Solar-Pro-3 and Qwen3-30B intervene frequently yet rank low on consensus gain.Intervention effectiveness aligns with consensus gain, and the three mediators tied at 24.6 hold the top three consensus-gain scores.
  • Socio-cognitive Adaptation Analysis: Every mediator contracts on at least one socio-cognitive axis, producing uneven capability profiles rather than a single frontier.GPT-5.4-mini and DeepSeek-V3.2 lose more under Multi-state Tracking than Gemini-3.1-FL and Qwen3-235B despite comparable overall consensus gain.
  • Strategy, Emotion, and Culture Shifts: Strategic posture causes the sharpest stress-test degradation, with Competing scores of 18.9–64.1 and Accommodating scores of 13.8–66.8.Qwen3-235B suffers the largest drops in both settings despite its high overall ranking.
  • Strategy, Emotion, and Culture Shifts: When both parties are reactive, every mediator drops, while cultural-distance effects are smaller but systematic.Emotional degradation does not follow model size, and scores decline as cultural distance from U.S. norms grows.
  • Strategy, Emotion, and Culture Shifts: The best intervention window is early for Strategy Adaptation or Emotional Regulation but later for Multi-state Tracking or Long-context Understanding.Timing profiles are compared over normalized 0–100% conversation progress across general and hard conditions.

6 Conclusion

SoCRATES benchmarks LLM mediators across eight domains and five socio-cognitive axes using automatic scenario construction and topic-localized evaluation. Its results show that conflict resolution remains challenging and varies across contexts and party compositions, making adaptation central to effective mediation.

  • SoCRATES probes LLM mediators across eight domains and five socio-cognitive axes.
  • The benchmark combines automatic scenario construction with a topic-localized evaluator.
  • Mediator performance shifts across contexts and party compositions, indicating that effective mediation hinges on adaptation rather than uniformity.

Limitations

SoCRATES currently evaluates mediation in English and focuses primarily on consensus. The authors identify multilingual mediation and broader measures of mediation quality as future extensions.

  • SoCRATES runs conversations in English even when parties receive different cultural identities, isolating cultural values from language variation but excluding multilingual mediation.Extending the benchmark could examine language choice, translation ambiguity, and language-specific politeness norms.
  • The benchmark prioritizes consensus because settlement outcomes are directly tied to consensus and can be scored consistently across domains.Party satisfaction, procedural fairness, trust restoration, and emotional repair are left for future evaluation with calibrated rubrics.

Ethical Considerations

SoCRATES uses simulated, anonymized conflicts rather than real disputants, while human annotators participate only in evaluator validation and persona-fidelity verification.

  • LLM agents role-play conflicts, so no real people participate as disputants in the benchmark process.
  • Scenario synthesis anonymizes residual references to specific individuals, organizations, or locations before scenarios enter the benchmark.Crowd-sourced and graduate annotators are recruited solely for evaluator validation and persona-fidelity verification.
  • The experiments use publicly accessible LLMs through APIs or Hugging Face checkpoints under their respective terms and licenses.Source scenarios are synthesized from deep-research seeds and do not incorporate text from any licensed corpus.

B Model Specifications

SoCRATES uses distinct LLM backbones for scenario search, writing, simulation, mediation, evaluation, and persona-fidelity validation across its experiments.

  • o4-mini-deep-research serves as Searcher for seed collection, while GPT-5.4 serves as Scenario Writer for recasting and condition-expansion rewrites.
  • DeepSeek-V3.2 serves as the Simulator for party-agent role-play and as the Evaluator for topic-localized scoring.
  • The benchmarked Mediators use the listed mediator pool, which also supplies Fidelity Simulators for persona-fidelity validation.

C Agentic Scenario Construction Details

The appendix details how SoCRATES constructs and expands conflict scenarios, simulates disputants, and validates persona fidelity and topic-localized evaluation. Its conditions vary one socio-cognitive factor at a time across a structured 15-condition testbed.

  • Scenario Construction: Scenario construction uses seed search, structured recasting, preference weighting, and persona-conditioned party simulation before benchmarking.The documented prompts cover seed collection, recasting, preference assignment, and role-playing simulation.
  • Condition Expansion: Each scenario has a general baseline plus 14 single-axis expansions, producing 15 conditions across five socio-cognitive axes.The axes include strategic posture, party composition, history length, emotional reactivity, and cultural identity.
  • Condition Expansion: Party-axis expansion adds one structurally distinct party while preserving the original parties and topics, increasing the number of states to track.The added party receives its own role, relation, and per-topic stances.
  • Condition Expansion: History expansion prepends four dated events, extending the background to roughly five times its default length.The original background remains appended as the final state.
  • Condition Expansion: Emotion and culture conditions alter party profiles through fixed reactivity endpoints and deterministic US, CN, or KR identity statements.Cultural identities summarize 0–100 Hofstede scores across six dimensions, while all parties interact in English.
  • Validation: Validation recruits qualified MTurk annotators for persona-fidelity comparisons and reports majority-vote agreement of α=0.75.Annotators compare dialogues differing only in the target party’s reactiveness level; fidelity is measured against the higher assigned level.

F.2 Consensus Alignment Annotations

Consensus alignment annotations compare supervised human scoring with SoCRATES and a per-turn judge at snippet and trajectory levels. The topic-localized evaluator follows active issue progression more closely than the per-turn alternative and remains aligned under a backbone substitution.

  • Validation: Evaluator validation compares SoCRATES, a non-expert annotator, and a per-turn LLM judge against the mean of two supervised annotators using Pearson correlation.The comparison is performed at both trajectory and outcome levels.
  • Annotation Protocol: Consensus is annotated at the snippet level using option-level positions and 1–5 agreement scores, carrying scores forward when an issue is not mentioned.Each snippet contains one back-and-forth exchange plus the background, topics, options, and preceding snippet.
  • Trajectory Alignment: SoCRATES tracks expert consensus curves closely, while ProMediate’s per-turn judge fluctuates because it scores inactive topics at every utterance.The topic-localized evaluator identifies turns where a topic is active or a party shifts position, then preserves the score otherwise.
  • Backbone Robustness: SoCRATES preserves strong alignment with expert judgments when DeepSeek-V3.2 is replaced by Qwen3-235B-A22B-Instruct.This tests whether evaluator reliability transfers across backbones rather than depending on one model.
  • Mediator Interaction: The mediator intervention process separates a temperature-0.6 decision about when to intervene from generation of the intervention utterance.When the decision is true, the generated utterance is inserted before the next party speaks.
  • Intervention Analysis: Solar-Pro-3 and Qwen3-30B intervene roughly twice as often and much earlier than top mediators, without improving intervention effectiveness or consensus gain.The result indicates that early, frequent intervention does not substitute for substantive contribution.

H.2 Benchmark Stability Analysis

The stability analysis tests whether evaluator choice, simulator choice, and stochasticity alter SoCRATES findings. Results support broadly stable mediator comparisons, while limiting some ablations to representative mediators or general conditions.

  • Evaluator Robustness: Evaluator-backbone substitution changes average metrics by −2.0, +3.9, and +0.6 points, while mediator rankings correlate at ρ=0.862 for effectiveness and ρ=0.786 for consensus gain.Timeliness shows weaker ranking agreement at ρ=0.406 because it is more sensitive to relevant-turn selection.
  • Evaluator Robustness: The alternative evaluator preserves mediator ordering despite weaker agreement on intervention timeliness.The timeliness discrepancy is linked to evaluator sensitivity to which trajectory turns are considered relevant.
  • Simulator Robustness: Simulator-backbone robustness replaces DeepSeek-V3.2 party agents with Qwen3-235B-A22B-Instruct across three mediators and all eight situations.The ablation uses 600 scenarios and tests whether socio-cognitive performance gaps persist rather than establishing a full mediator ranking.
  • Multi-run Robustness: Multi-run robustness repeats the mediator phase twice, yielding three independent runs whose variance reflects both mediator and disputant-simulator stochasticity.Because of simulation cost, this analysis is limited to general conditions.
  • Implementation Documentation: The appendix documents the prompts, backbone configurations, and annotation templates used across scenario search, simulation, mediation, and evaluation.These materials include backbone configurations, condition-expansion prompts, intervention prompts, and templates for fidelity and consensus annotation.
Loading 2606.05563v1…