Source-linked AI summary
Not All Agreement Counts as Corroboration: Provenance-Conserving Multi-View Fusion for Typed Action Admission in Human-Robot Collaboration
Zekai Jin, Hanrong Zhang, Yihong Tang, Fei Hu, Zhen Dong, Yi Shao
TL;DR
Embodied systems need to distinguish predictive agreement from separately countable evidence because repeated inference can multiply agreement without adding an evidential origin. PACT supplies provenance-conditioned meet-within/sum-across fusion and typed action admission. Across controlled and offline HRC evaluations, its evidence budgets respond to provenance reassignment as predicted, while its guarantees remain conditional on the supplied partition and assumptions.
Problem
Source-local values cannot identify whether agreeing outputs repeat one origin or come from separately countable sources, creating an evidence-countability problem for action decisions.
Method
PACT uses a supplied provenance partition, retains coordinatewise support within each unit, accumulates only across units, and applies typed admission after selection.
Results
Provenance-partition aggregation reduces ncsAURC by 0.0557 relative to singleton aggregation, and camera-grouped PACT admits 47 of 57 Qwen3-VL-32B reference-consistent candidates with no observed reference-inconsistent admission.
Takeaways & Limitations
Agreement constitutes corroboration only when provenance permits separate accumulation, separating computational multiplicity from evidential multiplicity in typed action authorization.
Takeaways & Limitations
PACT’s structural guarantees are conditional on the supplied provenance partition, whose correctness is not established by the method.
Abstract
from arXiv · showhide
For embodied systems, predictive agreement alone does not determine whether evidence warrants action; evidential origin matters. Repeated inference over one observation can multiply agreement without adding evidence, while source-local values do not reveal whether outputs have separately countable origins. PACT treats evidence countability as a relational variable for provenance-conserving fusion and typed action admission. A supplied provenance partition defines countable units. PACT retains coordinatewise support shared within each unit, accumulates only across units, and maps unmet release conditions to hold, confirm, or fallback. Under the stated assumptions, source-local values cannot identify countability; the coordinatewise meet is the greatest budget satisfying singleton fidelity and insertion non-amplification, with coarsening monotonicity and fixed-partition stability. Across 31,200 evaluations in 48 scene clusters, PACT attains a common-support normalized risk-coverage area (ncsAURC) of 0.0861. Excluding the constructed adversarial-consensus arm, provenance-partition aggregation reduces ncsAURC by 0.0557 relative to singleton aggregation, while the corroboration contrast vanishes. On complete-source records, native scores favor PACT, but a common posterior-peak score narrows its difference from nested Dirichlet and favors product fusion. Reassigning provenance over unchanged predictions moves evidence budgets as predicted. In offline human-robot collaboration, eightfold within-camera duplication leaves 720 typed responses per checkpoint unchanged; camera-grouped PACT admits 47 of 57 Qwen3-VL-32B reference-consistent candidates with no observed reference-inconsistent admission in 60 episodes. PACT separates computational from evidential multiplicity: agreement constitutes corroboration only when provenance permits separate accumulation.
1. Introduction
PACT frames evidence countability as a decision variable for embodied action admission: agreement from repeated inference may lack a separate evidential origin. It therefore combines provenance-conserving fusion with typed release responses.
- Motivation and contributions: Repeated inference over one observation can produce agreeing candidates without adding an evidential origin, so the admission policy may request confirmation.The outputs remain within one provenance component despite their agreement.
- Method: Typed admission preserves the reason for non-release through HOLD, CONFIRM, and FALLBACK rather than collapsing all failures into rejection.The typed response identifies the applicable release condition.
- Method: PACT retains coordinatewise support within each provenance component and accumulates evidence only across components licensed as separately countable.This separates provenance-conditioned support from score-based retention and release authorization.
- Results: Excluding the constructed adversarial-consensus arm, provenance-partition aggregation lowers ncsAURC by 0.0557 relative to singleton aggregation while the corroboration contrast vanishes.The evaluation separates conservation-law effects from empirical ordering and calibration effects.
- Results: In offline HRC, camera-grouped PACT admits 47 of 57 Qwen3-VL-32B candidates consistent with the reference action, with no reference-inconsistent admission observed.Within-camera copies leave the admission behavior unchanged under the stated conditions.
- Motivation and contributions: PACT identifies evidence countability as relational because source-local numerical attributes cannot distinguish repeated derivations from separately countable sources.A supplied provenance partition provides the missing relation between outputs and evidential origins.
2. Related Work
Related work regulates contribution strength, dependence, retention, candidate quality, or execution safety, but does not by itself determine which retained outputs count as separate evidence. PACT adds that instance-specific countability layer before selection and typed admission.
- Adjacent fusion and uncertainty methods: Reliability, uncertainty, and conflict regulate contribution strength, whereas countability determines which retained outputs may accumulate as separate evidential units.These are distinct questions within multi-view and embodied-system fusion.
- Dependence-aware fusion: Dependence-aware fusion addresses repeated or shared information as covariance, information-flow, or dependence-regime problems.These approaches motivate conservative combination when outputs agree but may overlap evidentially.
- Grouping and countability: Hierarchical fusion makes grouping part of aggregation, while PACT uses a supplied partition to specify countable units without asserting statistical independence.PACT retains a coordinatewise meet within components and adds commensurate budgets across components.
- Selective prediction and clarification: Selective prediction and conformal methods decide which predictions to retain or when clarification is required, leaving joint support among retained outputs unspecified.PACT resolves countability before score-based selection, then returns HOLD, CONFIRM, or FALLBACK through typed admission.
- Embodied action systems: Candidate-generation, verification, and safeguard methods address search, policy constraints, hazards, confidence, or intervention, but candidate and evaluator counts do not establish evidential countability.PACT connects candidate evidence to action authorization as a separate accounting layer.
3. PACT: Provenance-Conserving Fusion and Typed Action Admission
PACT constructs provenance-conditioned evidence budgets, accumulates them across supplied components, selects an action contract, and applies typed admission. Its guarantees depend on commensurate evidence units and the supplied provenance relation.
- Provenance-conserving fusion: Sources with overlapping provenance parents form one component, whose evidence is conserved by a coordinatewise meet before component budgets are added.Missing expected sources remain represented in the provenance graph rather than erasing their roles or connections.
- Evidence and provenance: PACT maps source-local representations to nonnegative evidence vectors while supplied parent sets determine evidential countability.The provenance partition is formed from overlapping parent sets and connected components.
- Identifiability: Value-only fusion cannot distinguish identical numerical attributes assigned either to one duplicated provenance component or to separate components.The distinction requires provenance parent sets in addition to reliability, conflict, availability, and validity attributes.
- Guarantees: The coordinatewise meet is the unique pointwise greatest rule satisfying singleton fidelity and insertion non-amplification.Its cumulative budget inherits coarsening monotonicity and fixed-partition stability under the stated assumptions.
- Guarantees: Adding a source to an existing component cannot increase cumulative evidence, while an exact duplicate preserving its provenance parent set leaves the budget invariant.Refining the partition may increase cumulative evidence, whereas coarsening cannot create it.
- Selection and typed admission: Typed admission applies after score-based selection and returns the response associated with the first failed release condition.The ordered predicates test command consistency, source validity, risk support, and component corroboration before admission.
4. Experimental Design
The experiments test PACT’s provenance accounting, typed admission, comparator behavior, and robustness to altered counting structures across simulation and offline HRC settings.
- Evaluation strategy: 31,200 evaluations across 48 scene clusters test evidence accounting, selective retention, and typed action admission under varied source conditions.The benchmark combines language, geometry, and path-risk opinions over five action contracts and includes availability, validity, timing, quality, conflict, and urgency perturbations.
- Evaluation strategy: PACT’s paired interventions isolate countability from numerical evidence through exact copies, false splits, near copies, coarsening, bridges, and regrouping.These interventions test within-component invariance, insertion behavior, fixed-partition stability, and component identity.
- Admission: The admission pipeline selects an action contract, orders instances by a selection score, and stops at the first failed ordered requirement.Admission includes source-support, quality, conflict, risk, and component-corroboration conditions, with complete components contributing binary incidence separately from meet-derived evidence magnitude.
- Comparators: The benchmark compares value-only, product, Dirichlet, cautious, pooling, and PACT-related fusion rules under shared opinions, source attributes, folds, and instances.A coordinatewise-maximum ablation separates exact-copy invariance from insertion non-amplification, while native and alternative scores distinguish candidate formation from score-induced ordering.
- Evaluation metrics: Risk–coverage evaluation uses ncsAURC over shared support, with lower values indicating lower retained error averaged across common coverage.The primary support is [0.10, 0.39], while repeated-fold, joint-holdout, and perturbation analyses use [0.10, 0.35].
- Offline HRC: The offline HRC study evaluates 720 target–time cases from 60 disjoint evaluation episodes, with candidate evidence and event proximity entering admission as complementary requirements.The episode is the sampling unit, and the study measures transfer to unseen episodes within the same six tasks.
5. Results: Conservation, Ordering, and Admission
PACT preserves evidence under valid provenance grouping, while false splitting inflates budgets and can reverse rankings. Its results show that admission and score choice materially shape selective performance.
- Conservation under copy and provenance interventions: Exact copies within a provenance component leave PACT’s fused result unchanged, whereas assigning them to distinct components changes decisions.Distinct-component duplication produced a 0.3228 flip rate and mean posterior l1 drift of 0.3961 across 93,600 comparisons.
- Conservation under copy and provenance interventions: The coordinatewise-maximum ablation violates insertion non-amplification, increasing component budgets in 96.4% of tested instances and raising ncsAURC from 0.6294 to 0.7918.The meet increased component budgets in 0% of those instances, while the ablation changed 16.2% of predicted contracts.
- Conservation under copy and provenance interventions: A single representative fails to recover the component meet, producing ncsAURC 0.8639 before admission versus 0.6294 for PACT fusion.After admission, the corresponding values are 0.1668 and 0.0861; at target coverage 0.13, the representative produced 246 wrong admissions while PACT produced none.
- Conservation under copy and provenance interventions: False splitting raises singleton evidence budgets and increasingly worsens Δsplit as multiplicity grows.Δsplit rises from 0.0243 at m=1 to 0.3676 at m=8 and 0.4338 at m=32 over support [0.10, 0.39].
- Conservation under copy and provenance interventions: Grouping composition changes selective ordering even with equal component counts, raising ncsAURC by 0.1988 or 0.2013 under two alternative groupings.Both paired cluster-bootstrap differences were positive across 2,000 comparisons over scenes.
- Fusion methods and selection scores before admission: Before admission, PACT ranks best with its prespecified score 1 − u, but alternative scores reverse its ordering relative to other methods.PACT ranks first with 1 − u and last with posterior peak, top-two margin, inverse normalized entropy, and pooled correctness.
- Fusion methods and selection scores under a common admission policy: PACT attains ncsAURC 0.0861 versus 0.1479 for nested Dirichlet under complete method–score–admission rules.The difference is −0.0618, with no observed wrong admissions at target coverage 0.13 for either method.
- Fusion methods and selection scores under a common admission policy: Under a common posterior-peak score, PACT’s advantage over nested Dirichlet narrows and product fusion becomes competitive.On 19,200 complete-source instances, the scores are 0.0956 for PACT, 0.0979 for nested Dirichlet, and 0.0909 for product fusion.
6. Sensitivity and Decision-Layer Latency
Joint holdout preserves PACT’s aggregate advantage across tested evidence-construction settings, but condition-specific reversals and missing-source conventions delimit that result. The decision layer is also low-latency on the evaluated benchmark.
- Joint holdout: The PACT–nested ordering persists when source-condition mixtures are retained and foldwise cutoffs are re-estimated, although joint holdout reveals condition-specific reversals.The aggregate ordering is therefore sensitive to held-condition shift rather than threshold refitting alone.
- Parameter sensitivity: Across 256 Sobol settings, post-admission ncsAURC[0.10,0.35] spans 0.0660–0.0691 for PACT, 0.1067–0.2231 for nested Dirichlet, and 0.1796–0.2864 for provenance-discounted pooling.PACT remains lower than both comparators in every tested setting and yields no observed wrong admission at the resulting 0.13 operating target.
- Latency: On the evaluated three-source CPU benchmark, PACT’s decision layer has median latency 195.8 μs and 95th-percentile latency 251.7 μs.The reported timing boundary and source-count scaling are specified in Appendix B.
7. External Evaluation on Learned Predictions and HRC Admission
External evaluations show that provenance reassignment changes evidence budgets and selective outcomes, while exact-copy invariance persists through typed admission. In HRC, complementary admission signals reduce reference-inconsistent admissions while retaining most reference-consistent candidates, but transfer and score dependence remain boundaries.
- Learned predictions: Across HandWritten and PIE dataset–model pairs, within-component copies leave PACT fusion invariant through multiplicity eight, while false splitting expands and all-view merging contracts evidence budgets.The largest expansion is 3.21× on PIE–RCML; the strongest contraction is 0.0016× on HandWritten–TMC.
- Learned predictions: False splitting expands Scene15–TMC’s mean evidence budget to 3.32×, whereas all-view merging retains 0.053× and both errors worsen accuracy and ncsAURC.Within-component copies leave accuracy at 98.0% for TMC and 98.5% for RCML; false splitting shifts posterior distributions by 0.21–0.25 in l1.
- Learned predictions: Better aggregate probability scores do not imply correct evidence accounting: false splitting sometimes lowers macro-averaged NLL, Brier score, and ECE10 while worsening ncsAURC.Across 12 HandWritten/Mfeat replicated-source choices, NLL and Brier reverse jointly in four choices for both TMC and RCML.
- Camera-acquisition grouping: At four prompts, acquisition grouping improves ncsAURC for Qwen3-VL-8B by 0.1202, while the Qwen3-VL-32B contrast includes zero and posterior confidence does not reproduce the 8B effect.The selective consequence of counting varies across tested checkpoints and scores.
- Camera-acquisition grouping: Acquisition grouping improves Qwen3-VL-32B accuracy by 3.16 percentage points but lowers Qwen3-VL-8B accuracy by 3.14 points relative to per-output counting.All-view merging lowers accuracy for both checkpoints, while selection-score ncsAURC changes only slightly in the reported comparisons.
- HRC admission: Target identity and event proximity together withhold 90 of 91 reference-inconsistent admissions while retaining 51 of 53 reference-consistent admissions allowed by candidate evidence alone.The inconsistent-admission rate falls from 12.64% to 0.14% over 720 repeated evaluations.
- HRC admission: The combined HRC rule retains 24 of 27 to 54 of 57 reference-consistent candidates and permits one reference-inconsistent admission per checkpoint despite recall ranging from 45.0% to 95.0%.For Qwen3-VL-32B and InternVL3-8B, inconsistent admissions fall by 11.81 and 9.72 percentage points, respectively.
- HRC admission: Within-camera duplication to multiplicity eight leaves all 720 final typed responses per checkpoint unchanged; camera-grouped PACT admits 47 of 57 Qwen3-VL-32B candidates with no observed inconsistent admission.For Qwen3-VL-8B, it admits 43 of 53 reference-consistent candidates, also with no observed inconsistent admission.
8. Discussion
PACT distinguishes provenance-based evidence counting from computational agreement, while separating evidence accumulation, corroboration, and typed admission. The discussion also identifies empirical and scope boundaries, including calibration, temporal evidence, and HRC evaluation limits.
- 8.1. Provenance partition as an evidence-counting variable: PACT adds provenance as a relational variable determining whether outputs may contribute support separately.This complements source reliability and conflict by representing evidential origin directly.
- 8.2. Implications for multimodal foundation models and embodied systems: Exact copies within one provenance component leave evidence unchanged when eligibility and remaining admission conditions are unchanged.This is tested through copy interventions and applies to offline HRC admission.
- 8.1. Provenance partition as an evidence-counting variable: 0.0557 ncsAURC reduction remains after excluding the adversarial-consensus arm, while the corroboration contrast disappears.The two operations therefore affect different stages of the decision path.
- 8.1. Provenance partition as an evidence-counting variable: Budget conservation is structural, whereas score ordering and probability quality remain empirical and can diverge under false splitting.Aggregate metrics therefore need not expose a miscounted evidential unit.
- 8.2. Implications for multimodal foundation models and embodied systems: Temporal within-episode changes are more transferable across tasks than absolute scores, but temporal comparison requires an earlier observation before release.The ordered policy returns CONFIRM for a target mismatch and HOLD otherwise.
- 8.2. Implications for multimodal foundation models and embodied systems: The HRC study measures reference consistency rather than physical success, and camera-grouped PACT recalls 0 of 10 reference-READY cases in the tested four-prompt setting.No observed reference-inconsistent admission does not establish physical safety or success.
9. Conclusion
PACT conserves evidence by retaining shared support within provenance components and accumulating only across separately countable components. Its empirical budget response is stable under provenance reassignment and within-acquisition duplication, while calibration and risk–coverage ordering remain empirical.
- 9. Conclusion: PACT's coordinatewise construction is the greatest budget satisfying singleton fidelity and insertion non-amplification under a supplied provenance partition.It also provides exact-copy invariance and coarsening monotonicity under the stated evidence assumptions.
- 9. Conclusion: Provenance reassignment over unchanged learned predictions moves evidence budgets as predicted across multi-view and multi-camera settings.Within-acquisition copies leave offline HRC admission unchanged when eligibility and other admission conditions remain unchanged.
- 9. Conclusion: Calibration and risk–coverage ordering remain empirical properties rather than consequences of budget conservation.Temporal admission imposes a separate constraint, with within-episode changes transferring more consistently than absolute score levels.
A. Proofs and Additional Stability Results
The appendix provides a symbol reference for the paper's main text and proofs.
- A. Proofs and Additional Stability Results: Table 11 summarizes the symbols used in the main text and appendix.It serves as a notation reference for the formal development.
A.2. Proofs of the main-text propositions
The appendix formalizes supporting-component counts, provenance-aware conservation, and the typed decision maps used by PACT. It also states implementation costs for parent-set indexing, overlap discovery, and sparse aggregation.
- A.2. Proofs of the main-text propositions: M_ev counts supporting provenance components once, while the component meet separately determines retained support magnitude.A non-supporting bridge remains in its component and cannot create another supporting component.
- A.2. Proofs of the main-text propositions: Provenance-aware conservation forbids cumulative evidence increases from adding sources to existing components and preserves evidence under exact duplicates.The definition also distinguishes sources that form disconnected components.
- A.2. Proofs of the main-text propositions: The typed admission map assigns HOLD, FALLBACK, CONFIRM, or admission according to command, structural, risk, and provenance-condition failures.The indicators f_cur, f_struct, f_risk, and f_prov identify the corresponding unmet conditions.
- A.2. Proofs of the main-text propositions: Exact within-component copies preserve the evidence vector, parent assignment, eligibility, and remaining conditions, leaving the final decision unchanged.This case map formalizes the decision path summarized in Algorithm 1.
- A.2. Proofs of the main-text propositions: With indexed parent sets, overlap discovery and aggregation admit sparse-graph reductions in time and space.The appendix gives bounds involving J, K, P_max, and the discovered-edge count.
A.2.2. Proofs
The proofs establish PACT’s meet-based evidence budget as the greatest rule satisfying singleton fidelity and insertion non-amplification, with stability guarantees for predictions and selection.
- Provenance-conserving budget: The coordinatewise meet is the unique pointwise greatest admissible rule under singleton fidelity and insertion non-amplification.The proof bounds any admissible rule by the meet and then shows the meet itself satisfies both requirements.
- Provenance-conserving budget: Summing coordinatewise meet budgets across provenance components preserves greatestness and prevents inserted sources from amplifying existing evidence.Exact copies remain in their original component and cannot merge previously distinct components.
- Partition stability: Coarsening a provenance partition cannot increase the resulting aggregate evidence budget.Each coarse-block meet is bounded by the nonnegative sum of its constituent fine-block meets.
- Prediction and score stability: With a fixed partition, posterior and selection score are Lipschitz in ΔΠ with global constants 2/W and 1/W, respectively.The global constants decrease as prior strength W increases, although the sharper score bound need not be monotone in W.
- Prediction and score stability: Margin conditions preserve both the predicted contract and select-withhold status, thereby preserving the selection-stage decision y(0).The posterior condition maintains top-versus-competitor gaps, while the score condition prevents threshold crossing.
- Evaluation setup: The benchmark combines language, handover geometry, and path risk into opinions over five action contracts, then compares PACT with alternative fusion methods.PACT uses the same concentration as nested Dirichlet under equal source scaling; score comparisons include posterior peak, top-two margin, inverse normalized entropy, and a regularized correctness model.
B.3. Learned-prediction evaluation details
The learned-prediction evaluation combines multiple multimodal checkpoints, tasks, episodes, temporal windows, cameras, and prompt variants while varying repeated computation and acquired viewpoints separately.
- Datasets: The evaluation uses TMC and RCML datasets for budget, posterior, accuracy, and ordering analyses, with Scene15–TMC extending full provenance and selective-ordering analysis.PIE–RCML supports budget analysis, while HandWritten/Mfeat supports additional TMC and RCML analyses.
- Evaluation scope: The evaluation spans six tasks with 120 episode-target and 600 counterfactual-target queries across six model families and four checkpoint pairs.The checkpoint panel includes Qwen3-VL, InternVL3, LLaVA-OneVision, and SmolVLM2 variants.
- Prediction construction: Event proximity is estimated with regularized logistic regression over frozen visual and text features using episode-grouped cross-validation.Primary and task-held-out models use distinct development-data configurations, with task-held-out models excluding the evaluated task.
- Evaluation scope: Each checkpoint supplies 2,400 outputs from 60 episodes, two temporal windows, five cameras, and four prompt variants.The design varies prompt multiplicity, camera removal, and exact-copy multiplicity to separate repeated computation from acquired viewpoints.
- Statistical testing: Temporal tests permute complete early-score profiles within task and event-occurrence strata, retaining episode structure in the concordance statistic.The main analysis removes the sole singleton block and uses 1,126 events in 695 episodes.
B.4. Sensitivity analysis and Scene15 evaluation
Sensitivity experiments vary language confidence, concentration, occlusion, quality, and conflict while assessing score scaling, provenance reassignment, and decision-layer latency on Scene15–TMC.
- Sensitivity analysis: The sensitivity design uses 256 scrambled Sobol configurations spanning language peak probability, opinion concentration, occlusion strength, quality, and conflict multipliers.Geometry and risk perturbations preserve the original top class, while fusion concentration remains fixed within each fold.
- Sensitivity analysis: 0.0861 is the equal-scale PACT ncsAURC after admission, compared with 0.1479 for nested Dirichlet under the reported scaled configuration.Training NLL selects PACT source scales of (0.63, 0.63, 2.52), yielding pre-admission values of 0.6323 versus 0.7749 and post-admission values of 0.1147 versus 0.1479.
- Scene15 evaluation: Scene15–TMC reports mean accuracy of 0.6679 across 30 stratified 80/20 train-test realizations, 0.95 percentage points below the published value.The provenance interventions operate on held-out predictions from five realizations without changing the learned outputs.
- Normalization sensitivity: 71.82% is the mean accuracy under the separate train-fitted normalization sensitivity across five realizations.The population standard deviation is 1.88 percentage points over 4,485 held-out predictions.
- Latency: Decision-layer latency ranges from 1.0 to 150.8 μs for provenance discovery on sparse graphs, with 95th percentiles from 1.1 to 166.4 μs.Timing excludes threshold fitting, provenance acquisition, perception, inference, communication, and physical execution.