Source-linked AI summary

COVER: Identifiable Evaluation of Coalition Routing

Raghul Sugumar, Amrit Gopinath

arXiv:2608.28475v1cs.AI

TL;DR

End-to-end gains cannot isolate coalition-routing quality because changing teams also changes messages and finalization. COVER fixes the information boundary, downstream stack, and finite legal team family, then uses complete or minimal-support interventions to identify the relevant contrasts. Across controlled and natural evaluations, it exposes measurable selection headroom while rejecting a universal routing-superiority claim.

  • Problem

    Changing a multi-agent team changes private messages and final-answer synthesis, so end-to-end accuracy alone does not identify coalition-selection quality.

  • Method

    COVER is an evaluation contract that fixes the public information boundary, downstream stack, finite legal team family, and coverage needed for absolute or relative identification.

  • Results

    0.768 declared-family oracle versus 0.637 prospectively frozen-router safe-evidence completion in ToolSandbox, while controlled tables establish exact selection evaluation.

  • Takeaways & Limitations

    COVER exposes selection headroom without manufacturing a routing win and separates routing effects from downstream synthesis within its declared scope.

  • Takeaways & Limitations

    COVER does not establish stack-invariant, large-pool, or general agent-system superiority, and one-draw provider execution is not population utility.

Abstract

from arXiv · show

When a multi-agent system changes its team, it also changes the messages and final answer it produces, so an end-to-end accuracy gap does not by itself identify a routing effect. We introduce method, an evaluation contract that fixes a public information boundary, downstream stack G, and finite legal team family before outcomes are generated. Complete coverage identifies exact finite-benchmark oracle regret conditional on that stack. For any finite collection of frozen policies, executing the union of their distinct selected teams is the minimal assumption-free support for every pairwise policy contrast, though not for absolute oracle regret. Two controlled tables with source-ID-disjoint splits test the instrument. On MuSiQue-12, a pre-specified privileged positive control improves regret from 0.532 to 0.402; a later public-interface control reaches 0.424 versus 0.554 but is retrospective. On HotpotQA-4, a pre-specified public direct scorer improves regret from 0.313 to 0.110. In fixed-stack Llama execution, verified route regret improves by 0.190, while the raw-answer gain is 0.010 with an interval crossing zero. A five-family ToolSandbox variant-shift validation exhaustively evaluates 16 declared teams on 14 untouched task variants (224/224 valid rows): the declared-family oracle reaches 0.768 safe-evidence completion, while the prospectively frozen router gets 0.637 (regret 0.131), failing the predeclared 0.10 criterion. A later retrospective comparator reaches 0.655, matching all-workers with 4.57 versus 5.00 workers on average. Thus COVER exposes selection headroom without manufacturing a routing win. A crossed-stack diagnostic shows absolute scores depend on G but finds no detectable router-by-finalizer interaction. COVER is an auditable measurement methodology, not a claim of stack-invariant or universal agent-routing superiority.

1 THE PROBLEM IN ONE PICTURE

COVER isolates coalition selection by fixing what the router can see and the downstream stack, addressing the confound that team changes also alter messages and finalization. It distinguishes exact complete-table oracle evaluation from minimal support for frozen-policy contrasts.

  • Changing workers can alter private messages and final-answer synthesis, so an end-to-end accuracy gap does not identify a routing effect.
  • COVER fixes a public information boundary and downstream stack G before evaluating legal teams.The stack includes workers, timing, message order, prompting, finalization, judging, and utility.
  • Complete observation of all legal teams makes finite-benchmark oracle regret a direct table lookup without modeling unseen coalitions.With missing teams, sampled-action regret may be reported, but exact oracle regret requires additional assumptions.
  • The route union is minimal assumption-free support for relative contrasts among frozen policies, whereas the complete family is required for absolute oracle regret.
  • COVER is a methodology whose routers serve as validation instruments rather than a universal architecture.

2 RELATED WORK: WHAT COVER IS AND IS NOT

COVER narrows routing evaluation to coalition selection under a declared finite family and common downstream protocol. It complements, rather than replaces, model-routing, team-scoring, off-policy, and learned-surrogate approaches.

  • Multi-agent routing can select roles, communication paths, or workers, whereas COVER targets coalition-selection quality under a common downstream protocol.
  • A 30-paper audit found that “routing” spans several estimands, including joint-system design, communication topology, sequential policies, and team selection.The audit was frozen, purposive, and non-systematic, with detailed limitations confined to Appendix F.
  • Set encoders, submodular selection, DPPs, and correlation clustering appear only as validation-policy ingredients, not as COVER’s claimed algorithmic contribution.
  • Unlike off-policy and surrogate methods that extrapolate to unobserved actions under assumptions, COVER completely evaluates a small declared action space.This makes the finite benchmark oracle observable rather than estimated under a surrogate model, at higher upfront intervention cost.
  • ToolLLM- and DyLAN-style comparators adapt published selection ideas to a frozen finite coalition table without claiming head-to-head replication.

3 THE EVALUATION CONTRACT

The COVER contract fixes the legal action space, information boundary, downstream protocol, coverage regime, and inference procedure. Its identification hierarchy separates exact absolute regret from assumption-free relative comparisons among frozen policies.

  • Evaluation contract: The action space declares legal teams before outcomes are generated.
  • Evaluation contract: The information boundary exposes task text and public cards while hiding evidence, answers, required-team labels, and outcomes.
  • Evaluation contract: The downstream protocol keeps workers, ordering, communication, utility, and finalization identical across routers.
  • Coverage and identification: Complete coverage identifies absolute oracle regret, while the declared route union identifies relative frozen-policy contrasts.
  • Inference: Inference freezes routers before held-out comparison and resamples tasks rather than correlated worker pairs for uncertainty.
  • Coverage and identification: Partial coverage cannot identify exact oracle regret in general, although structural or generative assumptions can support estimates when reported as such.
  • Coverage and identification: The route union is the unique inclusion-minimal assumption-free support for all pairwise contrasts, requiring exactly the number of distinct selected teams.

4 TWO CONTROLLED INTERVENTION TABLES

COVER evaluates coalition selection using complete controlled intervention tables with source-disjoint splits and a fixed protocol. The results distinguish sensitive validation instruments from deployable or confirmatory routing evidence.

  • Controlled tables: Exact regret is measured over two intervention-complete multi-hop QA tables, with source IDs disjoint across training, development, and held-out sets.The unit of inference is the task, and each held-out task has exactly one positive team: 300/300 HotpotQA tasks and 500/500 MuSiQue tasks.
  • Policies: UNIFIED-COVER scores proposed teams from task and public capability-card representations using a permutation-invariant frozen-encoder and MLP architecture.It is trained on observed training-coalition outcomes with regression, ranking, and best-team losses, without positional worker identifiers.
  • Results: MuSiQue-12 regret improves from 0.532 to 0.402 for the pre-specified privileged positive control, while the later public control reaches 0.424 versus 0.554 retrospectively.The privileged pair tests instrument sensitivity using latent component labels; the public-interface result cannot replace prospectively frozen validation.
  • Results: HotpotQA-4 compares public direct scorers in a four-action decision, testing inductive bias and sample efficiency rather than large-scale coalition-routing capability.Both baselines consume public task/card text and training-only best-full-team labels and can represent all four decisions.
  • Evidential status: Figure 2 separates retrospective MuSiQue comparisons from pre-specified HotpotQA comparisons; regret is exact within each declared action space, and scales are not compared across instances.This preserves the evidential status of the two controlled comparisons.

5 EXECUTION DECOMPOSITION UNDER A FIXED STACK

Under a fixed downstream stack, COVER decomposes execution into relative route comparisons supported by the selected-route union. The Llama study shows a clear verified-evidence contrast but only a small, uncertain raw-answer difference.

  • Fixed-stack execution: The selected-route union supports paired route differences under identical workers, ordering, communication, utility, and finalizer, without identifying an execution oracle.The oracle term cancels in the paired comparison, so the estimand is relative route quality.
  • Verified evidence: 0.190 executed-route-regret points is the verified evidence-complete Llama contrast, with 95% CI [0.140, 0.240] and Holm p = .00003.This confirmatory endpoint measures successful evidence/citation transport under the enforced verifier.
  • Raw answers: 0.010 is the untouched raw-answer contrast, with CI [−0.0067, 0.0267], so the interval crosses zero.The raw-answer result is conditional on the downstream configuration G rather than being the primary execution endpoint.
  • Crossed-stack diagnostic: Absolute execution values depend on the finalizer, while the crossed 2×2 diagnostic found no detectable router-by-finalizer interaction in its two-finalizer panel.All four within-cell route advantages were positive, but the broader six-configuration panel remains incomplete.

6 NATURAL HETEROGENEOUS-TOOL VALIDATION

The ToolSandbox study tests COVER in a natural heterogeneous-tool environment by exhaustively measuring safe-evidence completion across a finite coalition family. It finds substantial declared-family headroom, but the prospectively frozen router fails its performance criterion.

  • Study design: The declared ToolSandbox family contains 16 coalitions per task across five heterogeneous public tool families, with completion measured by official safe-tool-state milestones.The 16 teams are five singletons, ten pairs, and the full five-worker team.
  • Study design: All 224 expected held-out task–coalition rows are valid, and every task has at least two distinct coalition values.The validation uses 14 within-family task variants and passes both natural-gate prerequisites: tool use and coalition-dependent outcomes.
  • Results: 0.655 is achieved by a later retrospective learned value router, matching all-workers while reducing mean team size from 5.00 to 4.57 without improving value.Because the comparator suite was specified after held-out inspection, it cannot replace the prospective router result.
  • Interpretation: Exhaustive evaluation exposes 0.113 value left between the declared-family ceiling and the later best development-frozen comparator, separating measurement headroom from achieved routing performance.A final-answer-only benchmark would mix this gap with synthesis failures, while top-k success could hide it behind ties.

7 ARTIFACTS AND REPRODUCIBILITY

The paper releases artifacts that support deterministic re-scoring and recomputation of reported analyses, while documenting limits on replaying hosted executions and developing new public-information routers.

  • Released artifacts: The intervention archive contains hashed task identities, generic coalition identifiers, and deterministic values sufficient to re-score supplied route files.Source text and public cards are omitted for licensing reasons, so the archive cannot independently run a new public-information router.
  • Released artifacts: The ToolSandbox release includes development and held-out manifests, 544 development rows, 224 held-out rows, frozen routes, and retrospective comparator artifacts.The released summary, freeze record, and 84 per-policy held-out routes support recomputation of Table 6 without another model or API call.
  • Reproducibility limits: Hosted traces are not claimed to be perfectly replayable because providers may change despite temperature-zero requests.The paper releases compact outcomes, prompts, route decisions, model identifiers, protocol hashes, and no-API commands for calculations that do not require the provider.

8 DISCUSSION: WHEN SHOULD ONE USE COVER?

COVER is intended for exhaustive evaluation of a meaningful small team family under a fixed stack, with reporting that distinguishes finite-table, fixed-generation, and stochastic-execution quantities. Its oracle ceiling must not be mistaken for router performance.

  • Use cases: COVER is most useful when exhaustive outcomes are affordable for a small, meaningful team family and the researcher wants to isolate routing from downstream pipeline changes.It is particularly suited to comparing a selector against an established worker and finalizer stack.
  • Use cases: Theorem 1 shows that a precommitted route union is the minimal assumption-free support for all frozen-policy contrasts, but it leaves absolute oracle regret unidentified.Complete family coverage remains necessary for assumption-free absolute oracle regret.
  • Reporting: COVER recommends reporting whether a conclusion concerns a finite controlled selection table, fixed-generation evidence transport, or final task success under stochastic execution.These quantities are related but not interchangeable.
  • Reporting: 0.768 is the declared-family oracle ceiling, whereas the prospectively frozen router reaches 0.637 and the retrospective comparator reaches 0.655.Reporting the oracle without achieved performance would turn treatment variation into a fictitious routing result.

9 SCOPE AND CONCLUSION

COVER identifies routing effects only within a declared finite action space and frozen downstream stack, while separating relative policy contrasts from absolute oracle regret. The paper’s conclusion is methodological rather than a claim of universal routing superiority.

  • Scope: COVER asks how much benchmark value a router loses relative to the best team when only the selected coalition changes under a declared finite family and frozen stack.Complete interventions and a strict information boundary make this question identifiable.
  • Conclusion: COVER’s limits include construction-defined controlled outcomes, retrospective or limited natural validations, and one-draw provider execution that is not population utility.No result establishes stack-invariant, large-pool, or general agent-system superiority.
  • Scope: Moving downstream pipelines prevent routing contrasts from being identified, even with repeated observations of the two systems.The same observed outcomes can arise from a routing effect or from unchanged team values plus a downstream protocol change.
  • Scope: Complete coverage identifies team rankings and regret only conditional on one stack; it does not establish invariance under a different stack.A complete table under G0 can be compatible with either preserved or reversed rankings under G1.
  • Scope: A panel of stacks can summarize observed sensitivity, but its mean, variance, or pairwise agreement does not identify transport to an unobserved finalizer.The completed 2 × 2 study is a conditional diagnostic, while the broader six-configuration panel is incomplete.
  • Theory: Route-union support identifies relative frozen-policy contrasts, not absolute regret against the best team in the full legal family.The same support result extends to any finite number of frozen policies and deduplicates identical selected teams.

B BENCHMARK CONSTRUCTION AND INFORMATION CONTROLS

The benchmarks construct finite coalition-selection tasks with controlled public information, fixed downstream scoring, and source-disjoint splits. Their protocols distinguish privileged supervision and secondary controls from deployment-plausible public routing.

  • MuSiQue-12: MuSiQue-12 forms twelve-card pools by combining one target evidence component with three source-disjoint donor components, then permits any size-three coalition.The offline finalizer receives selected evidence in canonical order and scores recovered evidence units, creating a controlled complementarity test.
  • MuSiQue-12: The pre-specified Partition-COVER positive control uses train-only component-membership labels as privileged construction supervision, never as held-out routing input.A later public interface instead derives pair pseudo-labels from complete training coalition outcomes and uses public task/card text.
  • HotpotQA-4: HotpotQA-4 exposes four public worker cards and exhaustively evaluates the four possible size-three omissions against a fixed evidence-combining finalizer.This small action space tests sample efficiency and inductive bias rather than representation-theoretic separation from a classifier.
  • Information controls: The public-medium-33 protocol makes teams legal by deterministic hashed budget weights, while routers see only task text, card text, and weights.Its stopping rule, bootstrap, sign-flip test, and no-rerun condition were frozen before evaluation.
  • Scorers: Unified-COVER directly scores candidate teams using frozen sentence embeddings, task-conditioned card features, pooled set representations, and observed training-coalition outcomes.Partition-COVER instead combines public compatibility with exact balanced-partition completion using dynamic programming.

D.1 DEPENDENCE, REPRESENTATION, AND STACK CONTROLS

The secondary analyses probe dependence, representation, stack sensitivity, natural-tool provenance, and repeated execution. They preserve several positive route comparisons but explicitly separate retrospective, incomplete, or non-oracle evidence from confirmatory claims.

  • Dependence: A cluster-bootstrap sensitivity analysis retains a positive MuSiQue positive-control gain of 0.1300, with CI [0.0968,0.1634], across 201 source-component-sharing clusters.The analysis treats this as dependence sensitivity rather than evidence that every latent component is independent.
  • Representation: On HotpotQA, direct classifier regret is 0.4467, while distinguishable worker identity yields 0.1167 and constant worker profiles yield 0.7500.These post-hoc controls indicate that the hand-authored semantic coordinates do not explain the primary advantage, without overwriting the frozen 0.1100 result.
  • Stack controls: The completed 2 × 2 diagnostic finds absolute scores depend on the finalizer, while router-by-finalizer interactions include zero in both reported comparisons.The broader six-configuration panel remains incomplete, so no invariance claim follows.
  • Natural validation: The five-family ToolSandbox validation exhaustively evaluates 16 declared coalitions across 14 untouched within-family variants, with all 224 expected rows valid.Every task has at least two coalition values, passing the prerequisite gate that tools work and coalition changes affect outcomes.
  • Secondary reanalysis: At 12 cards, calibrated-graph regret is 0.3283 and its local-score gain is 0.0358, while COVER regrets are 0.1425, 0.1742, and 0.2533 at proxy budgets 6, 9, and 12.These secondary analyses broaden the controlled action space but do not establish natural large-agent deployment or cost efficiency.
  • Provenance: Post-Llama raw-output comparisons remain late prospective and are excluded from the main evidence spine despite being frozen before their own execution.The retained provenance is disclosed to avoid a file-drawer reading.
  • Repeated execution: Repeated execution on a 120-task subset yields raw accuracy 47.22% for Unified-COVER versus 40.28% for leave-one-out, with route-regret improvement 0.0694.The result supports repeat stability for this temperature-zero provider pipeline, not general population utility.
  • Reproducibility: The public archive supports rescoring supplied routes but excludes source-bearing questions, transcripts, provider metadata, and credentials needed to run new public-information policies.Researchers must obtain licensed datasets and reconstruct public inputs locally.

F FROZEN TARGETED ROUTING-ESTIMAND AUDIT

The targeted audit is a frozen, purposive snapshot of how recent papers use “routing” and distinguishes several intervention types. Its coding corpus and validation are deliberately non-systematic and single-coded.

  • Audit scope: The audit covers 30 recent papers purposively selected across team selection, communication topology, sequential routing, dynamic orchestration, and related areas.It is a targeted, non-systematic snapshot rather than a complete database screening.
  • Coding scheme: The coding unit is each paper’s declared primary intervention, classified as I, T, S, or J according to communication, team selection, sequential state-changing routing, or joint-system design.These categories distinguish different estimands rather than labeling other work invalid.
  • Audit limitations: The released audit corpus has a deterministic validator for record count, allowed codes, unique IDs, and reported-total consistency.The corpus was coded by one author without independent second coding, inter-rater reliability, complete candidate-screening counts, or database query logs.
Loading 2608.28475v1…