Source-linked AI summary

Epistemic Sybil Resistance: Multiplying AI Agents Without Multiplying Evidence

Marc Bara

arXiv:2609.01873v1cs.AIcs.MA

TL;DR

Multi-agent systems may multiply reports without multiplying evidence because reports can share hidden ancestry, while independent evidence can look similar. The paper formalizes this with conditional mutual information, derives shared-root information ceilings, and tests the claims with controlled language-model-agent experiments. The results support tracking evidential ancestry and dependence rather than agent, report, or representation multiplicity.

  • Problem

    Multi-agent orchestration makes report multiplicity an unreliable proxy for evidential multiplicity because reports may share evidence roots or model-induced dependencies.

  • Method

    The paper defines epistemic Sybil extensions using I(Θ; Z | R) = 0, analyzes Gaussian shared-root extraction with correlated errors, and evaluates the predictions in controlled agent experiments.

  • Results

    Theoretical and empirical results show that evidential ancestry and dependence, rather than report multiplicity or representation similarity, determine justified aggregation.

  • Takeaways & Limitations

    Collective inference should track conditionally novel evidence and evidential dependence instead of scaling confidence with apparent contributors.

  • Takeaways & Limitations

    The empirical study uses synthetic quarterly-revenue memos from one task family and one main model, so it makes no claim about scaling across model families.

Abstract

from arXiv · show

Multi-agent AI systems improve inference by spawning agents and synthesizing reports. But another agent is not another observation: apparently independent reports may descend from the same evidence, and genuinely independent evidence can produce nearly identical reports. We formalize this as an epistemic Sybil problem. A report Z is an epistemic Sybil extension relative to reports R when I(Theta; Z | R) = 0. No report-only aggregator can generally distinguish replication from independent corroboration: identical reports can warrant different posteriors under unobserved ancestry. A Gaussian shared-root model shows common ancestry does not imply complete redundancy. Repeated extraction adds information toward a source-level ceiling, and correlated extraction errors, which a shared base model can induce among independent agents, lower that ceiling further. We test these predictions with more than 20,000 controlled LLM-agent report and extraction calls on synthetic evidentiary documents. Holding one evidence root fixed while report multiplicity rises from 1 to 32 collapses naive posterior coverage from 0.940 to 0.263. Holding report count fixed while evidence-root multiplicity rises from 1 to 16 closes the gap, and the aggregators are statistically indistinguishable at k = 16. The agent's replicate extraction errors are correlated (gamma_cal = 0.719, estimated out of sample), and a correlated-extraction aggregator restores calibration accordingly. A controlled manipulation isolates representation similarity from evidential ancestry. It changes a report-space deduplication mechanism's mean inferred cluster count by 1.425 (95% CI [1.363, 1.485]), whereas a fourfold change in true ancestry changes it by only 0.040 ([-0.045, 0.120]). Collective inference should therefore track evidential ancestry and dependence, not agent or report multiplicity or similarity.

1 Introduction

Multi-agent systems can multiply reports without multiplying evidence: shared ancestry may look diverse, while independent evidence may look similar. The paper frames this as an epistemic Sybil problem and asks what information should govern aggregation.

  • Motivation: Generating reports is cheap, but one source can be repeatedly summarized, transformed, or retransmitted into apparently unrelated outputs.Conversely, genuinely independent evidence may produce nearly identical reports, so observable diversity and evidential diversity can diverge.
  • Problem: Agent multiplicity, report multiplicity, and evidence-root multiplicity are distinct quantities because orchestration can create dependence among reports.The system may reuse sources, retrieval results, prior outputs, or a shared base model.
  • Concept: An epistemic Sybil multiplies apparent evidential support without multiplying information about the latent state.The paper defines the relevant contribution relative to information already available to the aggregator.
  • Approach: The paper tests its analytical claims with a synthetic Monte Carlo study and more than 20,000 language-model-agent calls under controlled dependence.The empirical design also separates representation similarity from evidential ancestry.
  • Contributions: The paper introduces epistemic Sybil resistance as a criterion based on marginal conditional information rather than agent identity, report count, or report similarity.Its contributions include a report-only identification barrier and an explicit separation of agent, report, and evidence-root multiplicity.
  • Positioning: The work shifts collective inference from counting agreeing reports to assessing their evidential ancestry and conditionally novel information.This builds on established dependence concerns while targeting the architectural dependence created by generative orchestration.

3 Epistemic Sybil Resistance

The paper defines epistemic redundancy through conditional information and makes resistance relative to the aggregator’s information interface. Shared ancestry creates dependence but does not make additional reports automatically redundant.

  • Reports and roots: A report may come from a sensor, human, model, database, document, retrieval operation, algorithm, or composition of these sources.The framework treats all such outputs as reports received by an inference system about a latent state.
  • Reports and roots: Two reports can be conditionally independent given the state or derive from a shared evidence root, Θ → E → (X, Z).Common ancestry induces dependence but does not imply that one report adds no information after the other.
  • Definition: An additional report Z is an epistemic Sybil extension relative to admitted reports R when I(Θ; Z | R) = 0.A collection is treated analogously, and the definition concerns conditional information rather than surface duplication.
  • Definition: The incremental epistemic contribution Γ(Z; R) = I(Θ; Z | R) is representation-invariant but theoretical rather than directly observable.A Sybil extension need not duplicate an existing string, embedding, identity, argument, or source label.
  • Interfaces: Epistemic Sybil resistance depends on the information interface available to the aggregator, from report-only observation to authenticated provenance and complete joint structure.Richer interfaces can expose shared roots, derivation facts, or probabilistic dependence models.

4 No-Minting and Report-Only Non-Identifiability

The paper separates source-level information conservation from report-only identification. Descendants of fixed evidence cannot mint information, yet report content alone cannot generally distinguish replication from independent corroboration.

  • No-minting: Processing descendants of primitive evidence cannot create information about the latent state absent from their ancestry.The data-processing inequality and Blackwell ordering establish this source-level conservation principle.
  • No-minting: Repeated extraction need not be discarded because later descendants can expose additional information already latent in the source.Their total information remains bounded by the information in the evidential ancestry.
  • Non-identifiability: The same observable report profile can represent either cloned ancestry or independent corroboration, while the appropriate posterior differs.The binary construction uses identical observed reports under both structures but different conditional relationships.
  • Non-identifiability: No deterministic report-only rule with common marginal accuracy can be Bayes-correct under both clone and independent information structures.For p = 0.7, the unavoidable gap is approximately δ(0.7) ≈ 0.072, even when the rule knows p.
  • Implication: No rule under the trivial interface can guarantee both epistemic-Sybil invariance and Bayes-correct responsiveness to independent corroboration.Semantic similarity may still be informative, but it is not a universally valid identification criterion.

5 Information Aggregation under Shared Roots

Shared-root reports can reveal additional source information through repeated extraction, but their gains saturate at a source-level ceiling. Independent roots scale differently, and correlated extraction errors lower the ceiling further.

  • Scaling regimes: Exact replication adds no information: with ν^2 = 0, every report is the same observation and J_m = 1/σ^2.This is distinct from repeated extraction, where report-specific noise can be reduced.
  • Scaling regimes: Repeated extraction with ν^2 > 0 reduces extraction noise, but its information gain saturates because common source error ε cannot be averaged away.The marginal precision generated by later reports is positive but decreases with report count.
  • Scaling regimes: Independent evidence roots yield J_ind,m = m/(σ^2 + ν^2), so precision grows linearly with report count.Exact replication, shared-root extraction, and independent corroboration therefore have different information scaling.
  • Effective evidence: The effective sample size for correlated shared-root reports is m_eff = m/[1 + ρ(m − 1)], with discount κ_m = 1/[1 + ρ(m − 1)].Within the homoscedastic Gaussian equicorrelation model, this exactly matches the likelihood-precision correction.
  • Model boundary: The exact dependence correction is specific to the homoscedastic Gaussian equicorrelation model and is not generally Bayes-optimal for arbitrary evidence structures.Within the model, ρ is the proportion of conditional variance attributable to the common source component.
  • Correlated extraction: Shared model error creates an additional information ceiling that nominal agent multiplicity or prompt diversity cannot remove.The paper estimates this effect from replicate extractions by a real language-model agent.

6 Provenance as Side Information

Provenance can certify explicit derivation structure and zero-novelty extensions, but it cannot generally identify all dependence or guarantee that declared ancestry is truthful. In generative AI, certification must account for information in model parameters, not only retrieved documents.

  • Epistemic Sybil extensions: I(Θ; Z | R) = 0 defines an epistemic Sybil extension whose addition leaves the posterior unchanged.The posterior identity is P(Θ | R, Z) = P(Θ | R) almost surely.
  • What provenance can certify: Certified zero-novelty extensions preserve Bayesian posteriors when the interface is sufficiently informative, making dependence identification the difficult step.The invariant follows once zero conditional information is certified.
  • What provenance can certify: A provenance graph can expose authenticated derivation structure, but disjoint primitive roots do not establish conditional independence.Distinct datasets, shared environmental noise, or unrecorded common sources can still connect apparently separate reports.
  • Open provenance limits: Provenance may be falsified or withheld, so robust mechanisms must address truthful reporting and privacy-preserving structural certification.The paper identifies these as open problems rather than developing them further.
  • Model parameters as evidence: Model parameters can contain task-relevant information absent from retrieved reports, so paraphrase is not automatically a zero-novelty transformation.The paper therefore requires an independence assumption or a convention excluding parametric knowledge as admissible evidence.

7 Synthetic Model Validation

The synthetic validation implements the Gaussian shared-root model to test calibration under imposed dependence. Increasing reports from one shared root can make naive posteriors overconfident, whereas increasing root multiplicity adds information; these simulations validate the analysis rather than provide empirical evidence about real agents.

  • Experimental setup: The simulation varies report count n ∈ {1, 2, 4, 8, 16, 32} and root count k ≤ n using 60,000 Monte Carlo realizations per cell.The dependence-aware aggregator uses the correct shared-root covariance, while the naive aggregator treats reports as conditionally independent.
  • Report multiplicity: At n = 32 with one shared root, naive 95% coverage falls to 38.1% while the dependence-aware posterior stays near 95%.Naive negative log score rises from 1.22 to 7.27, while the dependence-aware score improves slightly.
  • Evidence-root multiplicity: At fixed n = 16, true information increases from 0.332 nats at k = 1 to 1.099 nats at k = 16, while the naive aggregator assigns 1.099 nats for every k.Shared-root precision approaches a finite ceiling, whereas independent-root precision grows linearly.
  • Interpretation: The validation imposes its dependence structure rather than measuring it, so its results are not empirical evidence that real agents follow the Gaussian model.The exercise benchmarks the analytical predictions before testing analogous behavior with real language-model agents.

8 Empirical Study with LLM Agents

Experiments with language-model agents show that report multiplicity can create overconfidence when evidence ancestry is fixed, whereas independent roots restore information. Correlated extraction errors and representation-sensitive deduplication further limit what report counts or similarity can establish.

  • Report multiplicity without evidence multiplicity: Naive 95% coverage falls from 0.940 at n = 1 to 0.263 at n = 32 with one evidence root, while provenance-aware coverage stays between 0.850 and 0.940.The naive calibration ratio rises from 0.961 to 5.473; provenance-aware calibration remains under 1.46.
  • Independent evidence roots restore information: The naive-provenance gap closes as independent root count rises, showing that root count—not report count alone—predicts the gap.At fixed n = 16, the aggregators are statistically indistinguishable at k = 16.
  • Correlated extraction errors and the information ceiling: Repeated extraction from one document adds limited information because block-mean variance declines from 9177 at m = 1 to 7510 at m = 32.The independent-extraction prediction falls from 7627 to 2678 and lies outside the bootstrap interval for every m ≥2.
  • Correlated extraction errors and the information ceiling: An out-of-sample estimate of correlated extraction, gamma_cal = 0.719, predicts the observed variance across all six extraction counts.The estimate uses 100 calibration worlds and disjoint evaluation data.
  • Correlated extraction errors and the information ceiling: The correlated-extraction aggregator restores Grid A coverage to 0.940–0.953 and calibration ratios to 0.95–0.96 across every report count.This extension is exploratory, while the plain provenance-aware aggregator drifts down to 0.850.
  • Representation similarity does not identify evidential ancestry: Representation manipulation changes inferred cluster count by 1.425, whereas a fourfold ancestry change shifts it by only 0.040.The representation effect has 95% CI [1.363, 1.485]; the ancestry effect has 95% CI [−0.045, 0.120].

9 Discussion

The discussion argues that collective inference should track evidential lineage and dependence rather than agent, report, or representation counts. Provenance helps expose structure that report content alone cannot, but it must be paired with an appropriate dependence model and remains subject to unresolved interface and manipulation problems.

  • Evidence multiplicity: Agent, document, model, vote, and report counts are not generally equivalent to independent evidence about a latent state.Reports sharing a root can retain partial information, while differently rooted reports can remain dependent through latent common causes.
  • Evidence multiplicity: Report multiplicity at fixed evidential ancestry can produce severe overconfidence when reports are treated as independent.This failure can arise when orchestrated agent or message outputs descend from shared evidence.
  • Representation and provenance: A report-space deduplication layer can track representation similarity more strongly than evidential ancestry, so similarity thresholds may not resolve the conflict.The discussion identifies this as a representation-tracking failure relevant to proposed defenses against shared-evidence overcounting.
  • Design implications: Reliable orchestration requires evidential lineage, overlap information, and dependence modeling rather than nominal or semantic multiplicity.Neither ancestry alone nor a dependence model alone is sufficient: ancestry does not specify the discount, while dependence requires something to condition on.
  • Design implications: When cross-correlation is unknown, robust fusion methods such as Covariance Intersection can preserve consistency for admissible correlations.The discussion contrasts this with independence assumptions that yield overconfident posteriors.
  • Open problems: Provenance interfaces face open problems because agents may suppress ancestry, coalitions may manufacture apparently separate roots, and both issues remain future work.The discussion notes that origin-bound authority and content-addressed lineage make such interfaces technically plausible, but does not resolve the problems.

10 Limitations

The framework’s scope is limited by empirical and theoretical constraints. The empirical study uses a narrow synthetic setting, while the general framework depends on distributions and provenance that do not capture every source of evidence or dependence.

  • Empirical scope: The empirical study is limited to one task family, one main model, and an artificial representation manipulation.These constraints bound how broadly the empirical findings should be generalized.
  • Framework limits: Conditional mutual information requires a probability distribution that is often unknown in real inference problems.This limits direct application of the framework when the relevant joint distribution cannot be specified.
  • Framework limits: Provenance records known derivation but not every latent common cause, including model parameters or human background knowledge that may constitute evidence.No additional document or sensor consultation does not by itself establish that no new information was used.
  • Framework limits: Provenance authenticity does not establish the truth of primitive evidence, and primitive evidence roots can themselves be Sybil-manipulated.The latter risk depends on whether their creation is constrained by domain-specific mechanisms.

11 Conclusion

The paper argues that collective inference should track conditionally novel evidence rather than agent, report, or representation multiplicity. It identifies provenance and dependence information as the key boundary for reliable aggregation.

  • I(Θ; Z | R) = 0 defines an epistemic Sybil extension whose additional report contributes no conditional information about the latent state.The definition is independent of agent identity and report representation.
  • Report-only inference cannot generally distinguish replicated ancestry from independent corroboration, even when observable reports are identical.Theorem 1 establishes this identification barrier, and the paper connects it to a concrete report-space failure.
  • Descendants of fixed evidence cannot collectively contain more information about the latent state than their evidential ancestors.The conclusion attributes this bound to the data-processing inequality.
  • The paper asks what minimum provenance or dependence interface would let an aggregator recover or approximate the posterior available under the full information-generating structure.This frames provenance, privacy-preserving provenance, and partially known dependence as connected design problems.
  • Collective inference should make marginal information its invariant because agentic orchestration can manufacture agents, reports, and representations at negligible cost.Nominal multiplicity therefore becomes an unreliable proxy for evidential multiplicity.

A Proof of Theorem 1

The proof uses identical observable inputs to show that deterministic report-only mechanisms must behave the same across distinct underlying information structures. Randomization cannot eliminate the resulting posterior-separation error.

  • Because the observable input is the same, a deterministic mechanism must choose the same value q in both structures.The structures nevertheless have different Bayes-optimal posterior values.
  • The deterministic result follows from the fact that identical observable reports require identical mechanism outputs despite differing underlying posteriors.The proof’s conclusion applies to the two information structures considered in Section 4.2.
  • E|Q − qC| + E|Q − qI| ≥ |qI − qC| implies that at least one randomized output error is at least half the posterior separation.This extends the identification barrier from deterministic to randomized outputs.

B Synthetic Model Validation: Additional Detail

The appendix provides the full metrics, tables, and figures for the simulation summarized in Section 7.

  • The appendix supplies the full metrics, tables, and figures for the Section 7 simulation.It expands the compressed presentation used in the main text.

B.1 Metrics

The validation evaluates calibration, uncertainty, scoring, and information accumulation under shared and independent evidence roots. Its results show that report count alone can misstate information and worsen scores under fixed ancestry, while dependence-aware aggregation accounts for repeated extraction.

  • Metrics: C = RMSE / reported posterior standard deviation, and C ≈ 1 indicates a correctly calibrated Gaussian posterior.Values substantially greater than one indicate overconfidence.
  • Metrics: E[∆ℓ] = I(Θ; Z | R) interprets expected realized log-score improvement as the conditional information contributed by an admitted report.Candidate aggregators are evaluated against this ideal rather than assumed to satisfy it.
  • Information accumulation: 1.099 nats is the naive aggregator’s implied information for every n = 16 case, while actual information ranges from 0.332 to 1.099 nats with independent-root count.The naive aggregator reacts to report count rather than evidential ancestry.
  • Shared-root evaluation: 1.22 to 7.27 is the rise in naive negative log score as report count increases from n = 1 to n = 32 under one fixed root.The dependence-aware Bayesian score instead improves slightly as repeated extraction removes part of the extraction noise.
  • Scope: The simulation is a constructed Gaussian-model validation, not empirical evidence that real LLM agents follow that model.Its narrower roles are implementation verification, demonstration of calibration errors from report-root mismatch, and hypothesis generation for later agent experiments.
Loading 2609.01873v1…