Source-linked AI summary

Exact Fusion and Coordinated Exploration in Multi-Robot Active Inference

Peng Wu, Mohsen Imani, Amidu Kamara, Md Tamzeed Islam, Seyede Fatemeh Ghoreishi, Mahdi Imani

arXiv:2609.17384v1cs.ROcs.MA

TL;DR

Robot teams can double-count shared beliefs during both posterior fusion and information-gain planning, causing redundant exploration. The paper corrects both stages with realized and expected evidence increments, and experiments show reduced redundancy and near-centralized value with linear-cost sequential coordination.

  • Problem

    Shared beliefs are counted repeatedly during fusion and planning, so the paper asks how teams can recover centralized posteriors and conditional joint information gains without sharing raw data or models.

  • Method

    The method adds realized evidence increments at fusion and expected increments from committed plans during coordination within conjugate exponential-family beliefs.

  • Results

    Across RockSample, foraging, and field monitoring, fusion correction alone leaves redundancy unchanged, while anticipated evidence removes it and sequential commitment recovers most centralized joint-planning value at linear cost.

  • Takeaways & Limitations

    A single increment-based message format supports exact belief fusion and coordinated exploration, with sequential commitment retaining the 1/2 greedy guarantee.

  • Takeaways & Limitations

    The theory assumes conjugate beliefs and conditionally independent observations, while the real-field experiment uses one interpolated time slice.

Abstract

from arXiv · show

Robot teams that learn a common environment model exchange belief summaries and plan by the expected information gain of their actions. Under conjugate exponential-family beliefs the shared belief is counted once per robot at two points: at fusion, the product of local posteriors counts the common prior $n$ times, and at planning, every robot scores its plan under the same belief and the team converges on the same unknown. Both errors are removed by adding evidence increments to the shared natural parameter, realized increments at fusion and expected increments at planning. The expected increment of a committed teammate gives the next robot its conditional gain; corrected gains sum to the joint gain, the redundancy removed equals the total correlation of the planned observation streams, and sequential commitment keeps the $1/2$ greedy guarantee. The expected increment is exact for Gaussian beliefs with fixed sampling paths and for Dirichlet beliefs under the novelty approximation of discrete active inference, whose team objective has a closed concave form within an explicit bound of the exact mutual information, and fails for finite hypothesis classes, where a short exact enumeration replaces it. Experiments on cooperative RockSample, foraging, and field monitoring show that fusion correction leaves exploration redundancy unchanged, anticipated evidence removes it, and sequential commitment recovers most of the value of centralized joint planning at cost linear in the team size.

I. INTRODUCTION

The paper addresses two forms of double counting in belief-sharing robot teams: fusion overcounts shared priors, while planning overcounts shared information gains. It corrects both by exchanging realized and expected evidence increments, with exactness conditions and empirical support across robot-team tasks.

  • I. INTRODUCTION: Fusion overcounts the common belief because every local posterior already contains the shared prior, whereas planning credits each robot with the same unresolved information gain.The planning overcount equals the total correlation of the planned observation streams and cannot be removed after execution.
  • I. INTRODUCTION: Adding evidence increments to the shared natural parameter corrects both fusion and planning errors while preserving a common environment model.Realized increments update fused beliefs; expected increments coordinate later plans.
  • I. INTRODUCTION: The expected increment of a committed robot supplies the next robot’s conditional gain, so corrected gains sum to the team’s joint gain.Sequential commitment retains the 1/2 guarantee of greedy submodular maximization.
  • I. INTRODUCTION: Expected increments are exact for Gaussian beliefs with deterministic paths and Dirichlet beliefs under the discrete active-inference novelty approximation, but fail for finite hypothesis classes.Finite hypothesis classes instead use short exact enumeration; the Dirichlet objective has a closed concave form within a bound on exact mutual information.
  • I. INTRODUCTION: Experiments show fusion correction alone leaves exploration redundancy unchanged, while anticipated evidence removes redundancy and sequential commitment recovers most centralized-planning value at linear cost.The approach targets model-parameter exploration rather than only agreement about current state, unlike related belief-sharing methods.

IV. EXACT FUSION BY EVIDENCE INCREMENTS

Exact fusion is obtained by retaining the common natural parameter once and adding each robot’s evidence increment. This recovers the centralized posterior and remains proper even when only subsets of increments arrive.

  • IV. EXACT FUSION BY EVIDENCE INCREMENTS: Naive fusion multiplies local posteriors and repeatedly counts the common prior and early evidence across synchronizations.In Gaussian and Dirichlet families, precision or total pseudo-count is multiplied by at least n at every synchronization.
  • IV. EXACT FUSION BY EVIDENCE INCREMENTS: The additive fusion update combines the shared parameter with each robot’s new evidence increment, recovering the centralized posterior exactly.The result follows from factorized likelihoods and conjugate exponential-family conditioning.
  • IV. EXACT FUSION BY EVIDENCE INCREMENTS: The additive information-form correction avoids tracking how many common-belief copies were received and remains proper under lost or delayed increments.Division with a stale common term can produce negative pseudo-counts or indefinite precisions.
  • IV. EXACT FUSION BY EVIDENCE INCREMENTS: For any subset of received increments, the additive parameter remains in the natural-parameter space and represents a proper posterior from the corresponding pooled data.This handles increments that arrive late or not at all without requiring division by a stale common belief.

V. COORDINATED EXPLORATION BY ANTICIPATED EVIDENCE

Anticipated evidence converts teammates’ future observations into conditional information gains, removing exploration redundancy while preserving sequential greedy guarantees. The approach is exact under a sufficient natural-parameter condition, but finite hypothesis classes require exact enumeration.

  • Coordinated exploration: Corrected conditional gains sum to the team’s joint information gain, eliminating redundancy equal to the total correlation of planned observation streams.Redundancy is nonnegative, vanishes for mutually independent predictive streams, and depends on plans rather than the later fusion rule.
  • Anticipated evidence: A committed robot broadcasts its expected evidence increment, allowing later robots to evaluate conditional gains under the anticipated shared belief.The same increment message has a realized form for fusion and an expected form for planning; exact conditional gains generally require teammates’ predictive models.
  • Approximation scope: The Dirichlet novelty approximation overestimates exact gain, with an explicit per-observation error bound that decreases as evidence accumulates.The relative error is O(v_j/N_j^2) when planned visit counts are random with variance v_j.
  • Sufficiency condition: The expected increment is sufficient when the gain depends only on an increment component fixed by the committed plans, yielding exact conditional gains after expectation.This condition supports Gaussian fixed-path sampling and deterministic-count Dirichlet novelty gains.
  • Finite hypothesis classes: For finite hypothesis classes, outcome-dependent likelihood increments invalidate the sufficient-message condition, so exact conditional gains require summing over the increment support.For m binary readings with common accuracy, the support has m + 1 points; at accuracy 0.9, the example returns 0.21 rather than 0.53 bits.
  • Greedy guarantee: Sequential commitment attains at least half the best assignment value when gains are evaluated exactly, including the full objective when its task-value term is monotone submodular.The guarantee applies to exact mutual information, the Dirichlet novelty objective, and compatible coordinated objectives.

VI. THREE CONJUGATE FAMILIES

The three conjugate-family cases instantiate anticipated evidence through Gaussian precision updates, Dirichlet visit counts, and finite-support enumeration. Gaussian updates are exact, Dirichlet planning has a closed concave objective near exact mutual information, and finite hypotheses use explicit averaging.

  • Dirichlet counts: Dirichlet anticipated evidence reduces to expected context visit counts, producing a telescoping concave objective that is monotone submodular.The corrected gain after teammate visits uses updated counts, and the resulting objective supports the same 1/2 greedy guarantee.
  • Dirichlet counts: The Dirichlet novelty objective remains within O(1/(N_j+n_j)) per planned observation of exact mutual information, with the error shrinking faster than the gain as evidence accumulates.The bound is uniform over teammates’ outcomes under deterministic visit counts.
  • Gaussian information form: Gaussian anticipated evidence is exact for deterministic sampling paths because observations increase precision independently of their values.The information gain depends on the Gaussian natural parameter only through precision, and the anticipated increment is the predicted information matrix.
  • Finite hypothesis classes: Finite hypothesis classes cannot use a sufficient mean increment because every likelihood-log increment depends on outcomes, so receivers evaluate a finite sum instead.For m planned readings, the binary common-accuracy case uses m + 1 support points per unknown.

VII. ALGORITHM

FEDAIF-AE combines exact fusion and coordinated planning in one synchronization loop by exchanging realized and anticipated evidence increments. It avoids transmitting raw observations or sender models, with fusion linear in team size and sequential commitment requiring short message rounds.

  • VII. ALGORITHM: FEDAIF-AE unifies fusion and planning corrections in one synchronization loop using realized and anticipated evidence increments.The algorithm forms each robot’s planning belief by adding earlier anticipated increments, then broadcasts its own anticipated and later realized increment.
  • VII. ALGORITHM: Each robot broadcasts sparse anticipated and realized vectors in shared-belief coordinates, so receivers need no observations, trajectories, or sender sensor and motion models.The exchanged vectors have dimension dim η, with optional dim ρ under Remark 1.
  • VII. ALGORITHM: Sequential commitment plans robots in a fixed order, evaluates candidate policies under the accumulated anticipated evidence, and broadcasts the committed plan’s expected increment.For finite hypothesis classes, the coordinated expected free energy uses the exact finite sum; Dirichlet planning additionally uses planned visit counts.
  • VII. ALGORITHM: During decentralized execution, each robot accumulates realized evidence increments and shares them after acting, allowing the common parameter to be updated without raw data.The realized increment is accumulated across the T execution steps and then used in the synchronization loop.
  • VII. ALGORITHM: Fusion costs O(n dim η), while planning adds one rollout per candidate policy and sequential commitment requires n short message rounds.Finite-hypothesis planning adds O(m) terms per unknown for the exact conditional-gain sum.

VIII. EXPERIMENTS

The experiments test the proposed corrections across RockSample, foraging, and field monitoring, manipulating fusion and planning independently. They compare team reward, posterior accuracy, and behavioral redundancy against naive, oracle, and local baselines.

  • VIII. EXPERIMENTS: The evaluation covers three testbeds matched to the theorem’s cases: finite-hypothesis RockSample, Dirichlet and binary foraging, and Gaussian field monitoring.Field monitoring uses both synthetic and real fields.
  • VIII. EXPERIMENTS: Fusion and planning are manipulated independently using naive or corrected fusion and naive or coordinated planning, with ORACLE and LOCAL reference conditions.The design tests whether each correction affects belief accuracy, reward, or exploration redundancy separately.
  • VIII. EXPERIMENTS: The evaluation measures team reward, divergence from the centralized posterior, and behavioral redundancy.Behavioral redundancy counts repeated robot-round sensing of the same target.

A. Cooperative RockSample

Cooperative RockSample evaluates finite-hypothesis coordination under randomized layouts and compares corrected fusion and exact or mean anticipated evidence with centralized and naive methods. Corrected fusion fixes belief accuracy, while coordination reduces co-sensing and improves reward, especially as team size grows.

  • A. Cooperative RockSample: Coordination with the exact gain removes 4.35 ± 0.58 co-sensing events and adds 2.80 ± 0.32 reward in RS(7, 8), n=3.Both effects are statistically significant, with p < 10^-70 for co-sensing and p ≈10^-54 for reward.
  • A. Cooperative RockSample: Corrected arms match the centralized posterior to 10^-12, whereas naive fusion yields DKL = 0.87–1.22.The comparison is reported for Table I’s RS(7, 8), n=3, 1,000 paired trials.
  • A. Cooperative RockSample: Fusion alone changes neither redundancy nor reward, while coordination removes redundancy and increases reward; the reward improvement replicates on RS(11, 11) and grows with team size.At n=5, the reward gap reaches +8.2 while naive co-sensing rises from 7.9 to 14.4.
  • A. Cooperative RockSample: The exact and mean anticipated-evidence arms are statistically indistinguishable, with a reward difference of +0.17 ± 0.26.The paper attributes the small difference to epistemic terms bounded by ln 2 nats per reading versus pragmatic stakes of ±10.
  • A. Cooperative RockSample: Under a misspecified common prior, fusion correction raises reward from 26.5 to 30.9 and belief accuracy from 0.72 to 0.95.Naive fusion repeatedly re-counts the wrong prior at synchronization.

B. Foraging: Dirichlet Yields and Centralized Joint Planning

The Dirichlet foraging variant gives robots unknown, non-consumable site yields and deterministic planned visit counts, satisfying the theorem’s exactness condition. CF-CP improves over naive planning, approaches centralized joint planning, and reproduces exact conditional gains numerically.

  • B. Foraging: Dirichlet Yields and Centralized Joint Planning: The foraging setup uses site yields in {0, 1, 2} apples with θs ∼Dir(1, 1, 1), and one commitment is one planned harvest.Three planners share model and preferences while differing only in coordination.
  • B. Foraging: Dirichlet Yields and Centralized Joint Planning: CF-CP earns 96, 189, and 323 versus 79, 123, and 187 for naive planning at n=2, 3, 4.At n=6, CF-CP earns 454 against 191 for naive.
  • B. Foraging: Dirichlet Yields and Centralized Joint Planning: CF-CP achieves 81–85% of the round-optimal objective while centralized joint planning earns 119, 257, and 442 at n=2, 3, 4.Sequential commitment therefore recovers much of the centralized planner’s objective at the tested team sizes.
  • B. Foraging: Dirichlet Yields and Centralized Joint Planning: The mean count increment reproduces the enumerated conditional gain to 10^-16 on every commitment.This matches the theorem’s prediction for Dirichlet beliefs with deterministic visit counts.

C. Multi-Robot Field Monitoring

Field monitoring shows that exact fusion and coordinated planning are both needed to avoid overconfidence and redundant exploration. CF-CP matches centralized conditional planning while preserving the exact centralized posterior.

  • CF-CP selects the same plans as sequential greedy conditional mutual information with trajectory sharing, differing only in its message format.The equivalence held on every seed by Theorem 1(a).
  • RMSE falls from 0.330 ± 0.014 to 0.129 ± 0.005 with coordination, with zero co-committed waypoints and performance 2.4× below LAWNMOWER.The paired improvement is −0.200 ± 0.014.
  • Naive fusion leaves RMSE at 0.727 ± 0.023 because multiplying prior precision makes predictive variance and information gain collapse.Even NF-CP reaches only 0.564 despite zero co-visits, since robots still plan under overconfident shared uncertainty.
  • Covariance intersection and Voronoi planning are within noise of CF-CP on model-matched fields, while exact fusion alone uniquely matches the centralized posterior.CI reaches 0.134 ± 0.005 and Voronoi 0.131 ± 0.003 versus CF-CP at 0.129 ± 0.005.

1) Real field:

On the Intel Berkeley temperature field, model mismatch changes the ranking: Voronoi planning outperforms information-driven coordination. Exact fusion with coordinated planning remains substantially better than naive baselines.

  • Voronoi planning achieves the lowest reported RMSE, 0.292 ± 0.002, ahead of CF-CP at 0.354 ± 0.007.The experiment uses one ten-minute reporting window, interpolated ground truth, and 100 paired seeds.
  • CF-CP improves over NF-CP, which reports RMSE 0.684, while NF-NP and CF-NP report 0.731 and 0.589 respectively.

IX. CONCLUSION

The paper concludes that additive evidence updates correct both fusion and planning errors, while experiments show their distinct roles and the limits imposed by model mismatch. The approach is efficient but depends on conjugacy, communication bookkeeping, and sequential coordination.

  • Adding realized evidence at fusion and expected evidence at planning removes shared-belief double counting, but exact fusion alone leaves exploration redundancy unchanged.
  • In field monitoring, centralized-planning mapping error is reached only when both corrections are used, while sequential commitment recovers most joint-planning value at linear cost.
  • CF-CP recovers half to two thirds of the gap between naive planning and centralized enumeration at linear cost and remains usable at n=6.Enumeration is infeasible at n=6.
  • On the misspecified real temperature field, Voronoi partitioning outperforms information-driven coordination.
  • The theory excludes learned belief representations and correlated sensor noise, while partial communication requires tracking which increments each robot has already added.Sequential commitment needs n message rounds; parallel group assignment weakens the guarantee.
Loading 2609.17384v1…