Source-linked AI summary
A Layered Analysis of Disagreement And Answer Quality in Multi-Agent LLM Debate
Chen Qian
TL;DR
The paper asks whether multi-agent debate elicits genuine, persistent disagreement and improves answers, rather than merely changing models’ reported or textual stances. It measures reported agreement, textual pushback, persistence after instruction removal, and open-weight token-logprob responses across committee debates and complementary tasks. Debate strongly changes what agents say, but the paper finds much weaker evidence that it changes persistent endorsement or improves final-answer quality.
Problem
Whether committee disagreement is genuine and survives scrutiny is rarely measured with condition-blind, reliability-checked instruments on open-ended tasks.
Method
The study separates reported agreement, textual pushback, persistence after stance-instruction removal, and open-weight stance responses across three-model committee debates and complementary evaluations.
Results
Debate changes expressed disagreement, but final-answer quality shows no detected gain: a swap-checked jury returns 299/299 ties, while verifiable accuracy is unchanged.
Takeaways & Limitations
Debate readily changes what agents say, whereas evidence that it changes what they persistently endorse or improves final answers is much weaker.
Takeaways & Limitations
The jury rules out only severe-magnitude quality differences, leaving moderate and smaller differences unresolved.
Abstract
from arXiv · showhide
Multi-agent debate, in which several LLMs exchange arguments before answering, is widely assumed to improve answer quality by surfacing genuine disagreement. That mechanism is rarely checked. We introduce four measurements: (A) the agreement a debater reports; (B) whether its reply text actually pushes back; (C) whether the position persists once the eliciting instruction is removed; and (D) for open-weight models, the stance response in the debater's own token log-probabilities. We evaluate three-model committees debating open-ended GlobalOpinionQA across 750 debates under three tones: friendly (seek common ground), neutral, and hostile (stress-test every position). (A) Tone strongly reshapes reported agreement: full agreement differs by 50.4 percentage points between the friendly and hostile endpoints. (B) A judge that reads only the reply text, never the self-report or the condition, recovers the same pattern. (C) The dissent appears partly tied to the instruction that elicited it: labels revert toward agreement 23.1 points more often after deleting the hostile instruction than under a matched re-ask that keeps it; question-weighted inference is inconclusive on first-round turns alone (p=0.0625), significant pooling all rounds (p=0.016), and only 11/28 first-round reversions also appear in the reply text. (D) Opposing arguments weaken a debater's stance margin more consistently than they shift its direction. For final answers we detect no quality gain: a bias-checked jury returns 299/299 ties (ruling out only large differences), accuracy on a verifiable control task is unchanged, and a jury without the bias check had declared debate the winner 66% of the time -- an artifact of reading order. Taken together, LLM debate readily changes what agents say, but we find much weaker evidence that it changes what they persistently endorse or improves the quality of the final answer.
1 Introduction
The paper tests whether debate produces genuine, persistent disagreement rather than merely changing what models report or say. Across four measurement layers, debate tone strongly changes reported and textual disagreement, while evidence for persistent belief change or improved final answers is weaker.
- Motivation: Multi-agent debate is assumed to surface genuine disagreement that improves answers, but whether disagreement is real and persists is rarely measured on open-ended tasks.The paper identifies self-reported agreement and debater-derived judgments as central measurement confounds.
- Four-layer framework: The framework separates reported agreement, textual pushback, persistence after instruction removal, and open-weight token-logprob stance responses.Layers A–C use the same committee runs, while Layer D uses a separate open-weight roster and proposition pool.
- Layer A: 50.4 percentage points: hostile instructions reduce full self-reported agreement from 74.8% to 24.4% relative to friendly instructions across 750 debates.The collapse is already present after one debate round and replicates the pilot pattern at five-times scale.
- Layer B: An external, condition-blind judge finds the same agreement collapse in reply text, with essentially no turns claiming pushback without textual pushback.When labels and text disagree, the judge reads the text as harsher than the label admits.
- Layer C: 23.1 percentage points: label reversion is higher after removing the stance instruction than under a retained-instruction re-ask, but first-round question-level evidence is inconclusive.Pooling all rounds is significant at p = 0.016, while only 11/28 first-round label reversions are confirmed in the text.
- Overall conclusion: The paper reports weaker evidence for persistent endorsement changes or answer-quality gains than for changes in what agents say.The bias-checked jury finds 299/299 ties, while the uncorrected 66% debate-win result reflects reading-order bias.
2 Related Work
Prior work spans debate, councils, judge reliability, and matched-compute comparisons, but no prior study combines open-ended deliberation with separated measures of textual disagreement, persistence, and position change. This paper presents its contribution as the intersection of these axes rather than as a wholly novel individual ingredient.
- Closest gaps: Matched-compute skepticism motivates comparing debate with sampling or self-consistency under equal budgets rather than unequal sample counts.The related work notes unequal budgets in prior comparisons and adopts matched-token discipline.
- Closest neighbors: Language Model Council uses non-deliberative voting, whereas ChatEval evaluates external answers and does not ask whether its debaters moved.These neighboring systems therefore differ from the paper’s position-change target.
- Closest neighbors: Peacemaker is closest to the persistence analysis but uses verifiable tasks and an unvalidated grader, while this paper studies open-ended items with judge validation.Peacemaker reports NAR and sycophancy measures, but the paper emphasizes differences in task setting and validation.
- Scope boundary: The paper withdraws an archived composition control because it pooled comparisons, lacked token matching, and lacked a judge outside both answer families.Consequently, that control supplies no composition conclusion.
- Closest gaps: Prior debate studies often evaluate verifiable-task accuracy or aggregate judgments without testing whether debaters change their own positions.The paper contrasts debate mechanisms, voting juries, and evaluator debates with its focus on position change.
- Paper position: The paper combines reported, textual, and persistent disagreement measurements with an external judge and matched-token baseline on open-ended deliberation.The authors describe the combination as the contribution, while calling the individual ingredients incremental.
3 Method: The A/B/C/D Framework and Instruments
The method separates four kinds of disagreement evidence: reported labels, textual pushback, persistence after instruction removal, and white-box stance responses. It validates the text judge with independence, known-answer discrimination, consistency, and a bounded human comparison, while using paired controls for persistence and logprob responses.
- Layers A–C: Layer A records each model’s structured agreement label, while an explicitness flag prevents schema defaults from appearing as disagreement.The label is generated in the same call that receives the stance instruction.
- Layers A–C: Layer B uses an external condition-blind judge to classify whether reply text challenges, adds evidence, concedes, refines, or restates.The judge sees only the parent turn and reply text, not the self-report or stance condition.
- Layers A–C: Layer C tests persistence by deleting the stance instruction and comparing label changes against a matched stance-retained re-query noise floor.The paired design accounts for temperature-1.0 label variability and reports excess change after removal rather than raw removal-arm change.
- Judge validation: The text judge must be outside graded model families, pass 80% known-answer discrimination for rebuttals and endorsements, and show study-data consistency.Gate A treats consistency and correctness as distinct requirements.
- Judge validation: Gate B compares judge–human with human–human agreement on 107 items; fine-severity difference is +0.03 with a 95% CI of [−0.09, +0.14], while binary-direction agreement is lower.The interval includes zero for fine severity but does not establish equality.
- Layer D: Layer D probes open-weight models’ next-token distributions after appending a fixed proposition-and-Likert elicitation, producing a continuous stance distribution without sampling.The probe renormalizes the seven option-letter logits and excludes degenerate reads under validity gates.
- Layer D: The logprob protocol contrasts opponent arguments with length-matched neutral filler, isolating paired TREAT−CONTROL effects across strong, weak, and irrelevant arguments.The contrast is signed toward the opponent’s pole to cancel shared context-growth, length, and format effects.
- Layer D: Conviction is measured by stance-margin erosion, whereas direction is measured by a heat-matched antisymmetric distribution shift toward the opponent.This separates weakened confidence from an actual change in preferred direction under saturation.
4 Experimental Design
The experiments use three-model committees debating open-ended survey questions under controlled debate tones, with scaled timing and logprob grids plus verifiable controls. The design isolates tone effects while tracking self-reports, blind textual judgments, persistence, stance probes, answer quality, and cost.
- Testbed: GlobalOpinionQA supplies cross-national opinion questions with multiple response options and no gold answer, so accuracy improvement cannot be claimed there.FRAMES and GPQA provide verifiable control tasks, while GPQA is also used as a contamination flag.
- Committee design: Each committee has three members, up to three debate iterations, shared routed history, and a chairman synthesis.Members respond to parent turns and emit both free-text opinions and structured self-reports.
- Tone conditions: The tone manipulation injects friendly, neutral, or hostile instructions only into debate-turn prompts, leaving initial opinions and aggregation neutral.The conditions respectively seek common ground, apply no added instruction, or stress-test every position.
- Tone conditions: The hostile condition is a Devil’s-Advocate manipulation that should mechanically affect Layer A, motivating independent Layer B and Layer C tests.The design distinguishes instruction-following from evidence about textual pushback or persistence.
- Scale: The pilot uses 10 questions per tone cell, while the timing grid uses 50 questions per cell across 15 cells for 750 debates.The larger grid varies when debate occurs across NONE, AFTER, and PER-TURN conditions and adds verifiable benchmarks.
- Measurement quality: The tone study records explicit labels on 99.7% of 1,189 pilot turns, reducing concern that differential defaulting explains the collapse.Only four labels defaulted in the blind-relabel analysis.
- Primary tone result: Pilot full-agreement rates are 69.9%, 50.9%, and 19.0% under friendly, neutral, and hostile tones, respectively.Table 2 reports these self-reported debate-turn shares for the three-model committee.
- Downstream evaluation: The answer-quality evaluation uses a bias-checked jury, while committee cost counts completion tokens including chairman synthesis.The jury design and token accounting support comparisons with matched-token self-consistency.
5 Results
Across the layered tests, debate reliably changes reported and textual disagreement, while evidence for persistent position change or improved answer quality is weaker. The results also show that judging safeguards and prompt-removal tests materially qualify apparent debate effects.
- LAYER A: −50.8pp separates friendly from hostile full-agreement rates in tri3, falling from 69.9% to 19.0%.The paired drop is robust under a record-level stance-permutation null (p = 0.0002).
- LAYER B: The blind judge reproduces the agreement collapse in reply text, with 81.1% exact self-report–judge agreement and essentially no unsupported pushback.The blind-over-self collapse ratio is 1.024, and the paired drop gap is −1.2pp with a confidence interval including zero.
- LAYER C: 23.1pp more first-round dissent labels revert after removing the hostile instruction than under retained-instruction re-asks.Question-level inference is inconclusive for first-round turns alone, while broader cuts reach significance; the blind judge confirms 11/28 first-round label reversions in text.
- Persistence and position: Models do not report much more mind-changing under hostility, while most mapped survey positions remain fixed across debate.Mean opinion_shift changes from 0.91 under friendly debate to 1.04 under hostile debate; cross-member entropy changes by −0.0157 bits.
- Answer quality: The reliable three-family jury returns 299/299 ties for debated versus undebated syntheses, reversing a first-pass 66% debate win rate.The quality jury’s mean swap-consistency is 0.89, above the ≥0.8 reliability floor, though its discrimination band excludes only severe differences.
6 Limitations
The paper’s limitations constrain the persistence, quality, and mechanism conclusions: several analyses are scoped to particular benchmarks, model rosters, protocols, or jury sensitivities. The authors therefore claim a confound-separated method and directional evidence rather than a settled empirical law.
- Scope and scale: The core instruments were developed on a small pilot, and the paper presents the method and observed direction rather than a settled empirical law.The timing grid replicates the Layer A/Layer B collapse at larger scale, but several analyses remain deliberately scoped.
- Layer C persistence: 23.1% attributable label reversion is based on 39 iteration-1 pairs, while corrected first-round analyses preserve the direction but remain inconclusive.The corrected question-equal rerun reports −11.8pp over 19 questions with p = 0.0625; the 9-question difference-in-differences reports p = 0.1250.
- Layer D generalization: The Layer D probe measures one scripted opponent exchange on a separate proposition pool and self-hosted open-weight roster, not the closed tri3 committee.It is a mechanism analog of a single debate exchange, so generalization to the committee setting is scoped.
- Benchmark and topology scope: The main mechanism study uses GlobalOpinionQA and one tri3 topology; separate web-harness results do not replicate the four-layer mechanism analysis.The paper also states that open-ended opinion items lack gold answers, so position change cannot be scored as improvement.
- Quality-jury sensitivity: The 299/299 jury ties exclude only severe-magnitude quality differences because the quality jury detects severe damage but not moderate or subtle damage.The jury’s mean swap-consistency is 0.89, but swap-consistency alone would also be passed by an always-tie judge; an authored discrimination battery was therefore required.
7 Conclusion
Across the study’s measurements, debate readily changes what agents say, but evidence that it changes persistent endorsement or final-answer quality is much weaker. Quality conclusions are bounded by jury sensitivity and task-specific scope.
- Disagreement measurement: 50.4pp separates friendly and hostile self-reported agreement in the three-round debate grid, with a 45.0pp gap after one round.The grid included 750 debates across a 15-cell timing design.
- Answer quality: 299/299 position-swapped jury comparisons tie, while verifiable accuracy remains timing-invariant.The jury’s swap-consistency is 0.89.
- Answer quality: A less reliable jury labeled debate the winner 66% of the time, but swap-checked judging erased that apparent advantage.The authors attribute the difference to reading-order bias.
- Limitations: The quality jury rules out only severe-magnitude differences because it catches severe synthesis damage at 100% but moderate damage at 0%.Moderate and smaller quality differences therefore remain unresolved.
- External task: +0.9pp accuracy on 235 paired web-research items falls inside the frozen ±7.5pp equivalence band.Two additional judges returned +0.4pp and +0.0pp, with 96.2% three-judge binary unanimity.
- Scope: The archived composition attempt supplies no Self-MoA conclusion because it pooled different comparisons, mismatched tokens, and lacked an outside judge.The study also makes no blanket claim that debate always costs more.
Ethics Statement
The study manipulates debate style rather than opinion content, using cooperative, neutral, or adversarial instructions within member debate signatures. Its committee roles separate independent opinions, debate turns, and final aggregation while preserving uncertainty and disagreement.
- Study scope: The manipulation targets cooperative versus adversarial debate style, not the content of GlobalOpinionQA opinions.The paper makes no normative claim about the survey propositions.
- Stance conditions: Neutral debate leaves the base debate signature unchanged.The neutral condition injects no additional stance instruction.
- Stance conditions: Cooperative debate instructs members to seek common ground, acknowledge reasonable points, and concede readily when warranted.Disagreement is retained only when the member has a strong, specific reason.
- Stance conditions: Adversarial debate instructs members to stress-test positions, identify weaknesses and counterexamples, and defend remaining disagreement plainly.The instruction rejects agreement merely for reaching consensus.
- Committee roles: Each committee uses independent initial opinions, per-round debate turns, and one chairman aggregation.All roles also carry a shared epistemic-humility clause requesting calibrated uncertainty.
- Committee roles: The chairman synthesizes final opinions, surfaces disagreement and its source, and weighs reasoning quality rather than assertion confidence.The aggregation should not redo the task or introduce new criteria.
C Committee implementation
The committee is implemented with typed DSPy signatures that preserve both free-text replies and structured self-reports while routing threaded debate turns. Timing is controlled by one program dial, with a serving caveat across API surfaces.
- Architecture: Typed DSPy signatures capture free-text opinions and structured self-reports in the same member call.Debate history is threaded through a specific parent turn and earlier context.
- Structured outputs: Each debate turn emits an agreement label and an opinion-shift score from 0 to 4 relative to the member’s previous turn.The agreement labels range from fully agreed to fully disagreed.
- Validation: A validator flags omitted agreement labels before defaults apply, preventing default inflation from masquerading as agreement or disagreement.Layer A analyses use only explicitly labeled turns.
- Timing modes: PER-TURN runs three debate rounds, AFTER runs exactly one round after independent drafts, and NONE sends drafts directly to the chairman.The stance instruction is injected only into the debate signature.
- Serving: Completion-token accounting uses raw generation history, including chairman synthesis, rather than a usage dashboard that undercounts member generations by approximately 3×.The two API surfaces may differ in hidden-reasoning token accounting.
D Sample debate transcript
A hostile-stance transcript shows members broadly rejecting a gender-role statement while debating the calibration between disagreement and strong disagreement. The external judge evaluates reply content independently of labels and conditions.
- Task and setup: The hostile example concerns whether men should earn money and women should care for home and family.Members respond on a four-option agreement scale.
- Initial opinions: All three members reject the statement, but distinguish sex-based duties from freely chosen household arrangements.Their reasons emphasize autonomy, equal opportunity, and non-coercive choice.
- Debate round: The first debate round preserves the direction of disagreement while contesting whether the proper calibration is disagree or strongly disagree.Both illustrated replies are labeled partially agreed with shift=1.
- Consensus: The chairman identifies autonomy and equal opportunity as the durable shared ground behind the convergence.The synthesis treats the statement’s roles as sex-based duties rather than freely allocated responsibilities.
- External judging: The judge receives only the parent argument and reply text, excluding self-reports, stance conditions, model names, and member identities.This isolates textual behavior from the eliciting metadata.
- External judging: The judge separately scores stance and move type, allowing agreement with a new argument or disagreement with mere restatement.Move types include new argument, refinement, concession, restatement, and challenge.
- Judge validation: Judges must achieve at least 80% recall for both rebuttals and endorsements on a balanced authored battery.The two-sided floor prevents an always-agree judge from passing through skewed pilot data.
F Gate-B human annotation protocol
The annotation gate used blinded, independent human labeling of debate turns, with self-reported agreement removed so text-based judgments could be evaluated separately.
- 173 debate turns were independently labeled by two annotators using a purpose-built web application.The seed was class-balanced and drawn from archived pilot transcripts.
- Annotators saw only the question, parent argument, and reply text, while the model’s self-reported label was stripped before rendering.This made annotation blind to Layer A by construction.
- Each item elicited an agree/disagree stance direction and a four-level agreed_level label.
G The pairwise quality jury
The pairwise quality jury compares debated and no-debate syntheses with order-swapped, position-consistent judgments, anchored sub-scores, and a sensitivity battery.
- Each question produced a debated synthesis and a same-question, same-stance no-debate synthesis for jury comparison.The protocol included AFTER or PER-TURN debate conditions versus NONE.
- A three-family jury scored every answer pair twice with the answer positions swapped, accepting a verdict only when the winner remained consistent.The pair’s verdict was the jury majority.
- Four anchored sub-scores—coverage, steelmanning, commitment, and soundness—were assigned before deriving the winner.Each criterion was scored from 1–5 separately for each answer.
- A 54-pair discrimination battery tested whether the jury could detect controlled quality degradations across three severity tiers.The battery was designed to complement the stance judge’s discrimination floor.
- The grid contained 15 cells of 50 questions spanning GlobalOpinionQA, FRAMES, and GPQA across specified stance and timing conditions.Nine cells covered open-ended GlobalOpinionQA; verifiable arms used FRAMES and GPQA under neutral stance.
I Corrected stated-position convergence analysis
The corrected convergence analysis preserves each question’s support structure and finds a small debate-arm entropy reduction, while the human-distribution contrast remains unresolved.
- The corrected analysis pairs neutral per-turn and no-debate arms on the same 50 GlobalOpinionQA questions while preserving 2–7 support options and a refusal category.Gemini mapped 450 unique opinions across three randomized option orders with a 100% parse rate.
- −0.0157 bits was the debate arm’s within-question entropy change, with a 95% question-bootstrap CI of [−0.0287, −0.0033] and sign-flip p = 0.0194.The no-debate arm changed by +0.0000 bits, so the primary difference in changes equaled the debate-arm change.
- The secondary final-to-final entropy contrast was −0.0104, but remained unresolved after Holm correction with p = 0.2376.
- A matched final elicitation yielded mean Jensen–Shannon distances of 0.120 for debate and 0.116 for no debate, with paired difference +0.0037 and Holm p = 1.0000.The paper therefore makes no accuracy, improvement, or representativeness claim from this comparison.
- Table 7 treats D−I as the primary comparison and permits only the preregistered Gemini row to license the frozen ±7.5pp equivalence decision.Luna and Sol are robustness evidence rather than replacements for the primary row.
J Supplementary web-research accuracy extension
The supplementary web-research extension compares isolated and cross-critiqued research workflows under a fixed benchmark setup, with independent judging and robustness audits.
- The extension uses a fixed LiveBrowseComp revision with 235 locally held-out money items selected by a frozen qualification-set power rule.It is not an official LiveBrowseComp reproduction, and the items are not called contaminated.
- S uses one Opus-5 agent, I adds an independent GPT-5.6 Luna researcher without draft visibility, and D adds one cross-critique revision before the same chairman.The I/D chairman, system instruction, member order, output contract, and token ceiling were word-identical.
- Both audit stages judge semantic correctness against the frozen benchmark reference using only the question, reference, and candidate answer.The preregistered Gemini judge evaluated all 705 arm assignments, while a frozen GPT-5.6 Luna pass tested robustness.
- The AI-only audit reported 97.0% Gemini–Luna, 97.6% Gemini–Sol, and 97.7% Luna–Sol binary agreement across judge pairs.All three judges agreed on 96.2% of the 705 assignments; the predeclared human spot check was not completed.
K Logprob instrument implementation
The Layer D instrument elicits stance distributions from open-weight models with a seven-letter probe, comparing opponent-context treatment against matched neutral control. Its design stratifies propositions by prior stance, controls for argument relevance and conviction, and averages repeated cells at the proposition level.
- Probe construction: The probe reads logits for seven stance letters from A (strongly disagree) to G (strongly agree) after appending a proposition-specific user prompt.Letter mass is renormalized over the seven options, with validity gates and a degenerate-read guard.
- Cell protocol: Each cell compares the receiver’s initial prior with a treatment containing an opposing argument and a length-matched neutral filler control.Arguments vary by strength and relevance, while the opponent’s pole is chosen opposite the receiver’s round-0 lean.
- Sampling and stratification: Propositions are stratified per receiver by prior stance into MID and POLE groups before the debate grid is run.The realized pool contains 27 of 44 authored bipolar public-policy propositions.
- Sampling and stratification: 612 cells were realized, with no probe reads excluded by validity gates; inference averages repeated cells within each proposition and weights propositions equally.The planned 768-cell grid was short because sparse prior strata could not fill every bucket, not because of probe-gate attrition.
- Instrument validation: The implementation subtracts pole-anchored heating effects so centered stance movement remains distinguishable from purely heated responses.A self-test requires heated pole changes to leave an approximately zero heat-free residual while preserving genuine centered location shifts.
L Reproducibility of numbers and figures
The paper automates numerical and figure generation from archived analysis artifacts and provides code and manifests for inspecting the execution and analysis path. Some benchmark materials remain excluded from the public release under a canary policy.
- Automated numerical consistency: Every quantitative value is generated from archived analysis artifacts by a build script rather than typed manually.Figure generation recomputes plotted statistics from raw records and asserts equality with the corresponding numerical macros.
- Code artifacts: The repository includes the committee engine, experiment entry points, relabeling and persistence analyses, and proposition-level logprob reduction code.These components expose the core implementation behind the paper-facing aggregates.
- Code artifacts: The supplementary web extension includes fail-closed coordination, shared protocol validation, isolated search tooling, blinded grading, and paired equivalence inference.The listed components implement and validate the shared web-search and evaluation stages.
- Release boundary: LiveBrowseComp plaintext, private trajectories, candidate answers, and judge rationales are excluded from the online release under the benchmark’s canary policy.The public release instead includes pinned dataset metadata, prompts, contracts, amendments, inferential rules, hashes, and aggregate counts.