Source-linked AI summary
Record Grouping Controls Evidence Weight in Language Models
Zhongxuan Liu, Sicheng Zhou, Hongzhi Wang
TL;DR
Retrieval records can expose repeated presentations as separate evidential contributions, raising the question of which information a pre-generation representation should preserve. The paper defines a partitioned group–content state that removes within-group copies while retaining complementary content and derives bounds for partition error. Across controlled experiments, false splits increased attack-side probability by 10.27–32.66 points, false merges reduced influence by 9.13–31.79 points, and campaign effects varied by checkpoint and presentation order.
Problem
Retrieval-record multiplicity can diverge from underlying evidence sources, leaving unresolved how to preserve complementary evidence while removing within-unit copies.
Method
The paper characterizes a supplied-partition group–content representation, proves invariance and equal-count separation, derives a content-aware error bound, and tests partition interventions with matched controls.
Results
False splits added 10.27–32.66 percentage points and false merges removed 9.13–31.79 points; controlled campaign effects were checkpoint-dependent and showed order interactions.
Takeaways & Limitations
The supplied partition is a controllable pre-generation representation variable that allocates model-facing evidential influence, while behavioral direction remains checkpoint-dependent.
Takeaways & Limitations
The analysis assumes claim-relative evidence roots and uses externally supplied operational grouping metadata; the data are public and include no new human-subject interactions.
Abstract
from arXiv · showhide
Retrieved records are presentation units; a supplied partition determines which records enter a language model as one evidential contribution. We characterize the invariant group-content state that removes within-group copies while retaining complementary canonical content, show that equal group counts can encode different evidence states, and derive a sharp content-aware partition-error bound. Given a supplied partition, our pre-generation representation deduplicates and aggregates content within groups and bounds each group's contribution. Across 104,402 trials and 6 public checkpoints, a central natural-text intervention finds that content-fixed false splits add 10.27-32.66 percentage points and false merges remove 9.13-31.79 points; a matched six-slot control retains the positive direction in all 16 cells. In a new 48-item controlled campaign panel, changing the supplied partition produces measurable, checkpoint-dependent decision shifts across all four models, and the balanced mirror design exposes substantial order interactions. Together, the theory and experiments establish the supplied partition as a controllable pre-generation representation variable and characterize its checkpoint-dependent behavioral effects.
1 INTRODUCTION
Retrieval records can multiply presentation opportunities beyond underlying evidence sources, making the supplied partition an evidence-weighting decision. The paper introduces a group–content representation that removes within-group copies, preserves complementary content, and supports controlled tests of false splits and merges.
- Packaging records can outnumber underlying evidence sources, so chunking, syndication handling, and record emission implicitly choose an evidence-weighting rule.
- Group count alone is insufficient because equal-sized partitions can combine different evidence elements and encode different evidence states.
- The experiments hold ordered evidence words fixed while changing the emitted partition, with a matched six-slot control fixing repeated non-content fields and tokenizer length.
- The representation removes exact copies, aggregates complementary passages, and assigns one bounded contribution per supplied group.
- The paper characterizes invariant and faithful group–content representations, proves equal-count separation, and derives a content-aware partition-error bound.
- A controlled panel extends the intervention to four cross-genre renderings of one campaign root and a four-independent-root control, revealing checkpoint-specific reversals.
2 RELATED WORK
Prior work shows that frequency, source cues, redundancy, and provenance affect language-model decisions and motivates approximate repetition-invariance methods. This paper instead studies how an external claim-relative representation makes replication invariance a property of the model-facing state before generation.
- Prior studies establish that contextual frequency, source cues, metadata, and authority can alter language-model decisions.
- The paper's novelty is making replication invariance a property of the model-facing state with enforcement independent of checkpoint responses.
- Existing work examines repeated or paraphrased support, credibility shifts, explanation effects, and approximate repetition invariance across language-model settings.
- Adversarial retrieval work addresses coordinated publication and provenance-aware aggregation, while this paper focuses on the state supplied before generation.
3 PARTITIONED EVIDENCE BEFORE GENERATION
The paper defines claim-relative evidence roots and a supplied partition as the pre-generation structure governing grouped content. Its quotient representation removes nuisance duplication and relabeling while preserving group–content distinctions, and its bounds connect partition changes to decision stability.
- 3.1 CLAIM-RELATIVE EVIDENCE ROOTS: Each claim-evidence atom has one claim-relative information-generating root, and independence is defined by distinct roots under that provenance relation.
- 3.1 CLAIM-RELATIVE EVIDENCE ROOTS: Block count represents epistemic multiplicity, whereas returned-record count represents presentation multiplicity; the distinction is claim-relative.
- 3.1 CLAIM-RELATIVE EVIDENCE ROOTS: Universal exact text-only recovery exists precisely when worlds with identical visible text share the same oracle partition.
- 3.2 INVARIANT AND FAITHFUL GROUP–CONTENT STATE: The canonical evidence quotient preserves distinct groups with identical content while deleting within-group canonical duplicates and ignoring record order and opaque-label renaming.
- 3.2 INVARIANT AND FAITHFUL GROUP–CONTENT STATE: Faithful invariant representations are exactly injective recodings of the canonical group–content quotient.
- 3.2 INVARIANT AND FAITHFUL GROUP–CONTENT STATE: Equal group counts and block-size multisets can still yield distinct bounded evidence states, with a scalar aggregator attaining EI(PA) = 2B and EI(PB) = −2B.
- 3.2 INVARIANT AND FAITHFUL GROUP–CONTENT STATE: Same-root proliferation leaves the number of outer contributions fixed and changes the state only through genuinely new canonical content.
- 3.3 FROM PARTITION ERROR TO DECISION STABILITY: The content-aware discrepancy cancels matching group–content sets, and its bound links changed overlap components to decision-margin stability.
4 EXPERIMENTAL PROGRAM
The experimental program evaluates how supplied grouping and downstream record partitioning affect model-facing evidence weighting while controlling candidate scoring, dependence structure, and serialization. It combines natural-text audits with a controlled campaign and independent-root panel.
- Panels and data: The study analyzes 101 PERSPECTRUM claims and 138 ConflictingQA questions as separate domains, with grouping panels using 66 HUMAN claims and all 138 WEB questions.
- Scoring: Candidate-choice audits compare two length-matched assistant answers using total conditional log likelihood, while controlled generation scores parsed leading Yes/No outputs.
- Scoring: The primary outcome is the attack-side candidate’s likelihood share, pattack, computed from the two candidates’ log likelihoods.
- Statistics: The analysis averages mirrored assignments and orders within items and uses fixed-seed 10,000-replicate bootstraps over items or natural-text dependence components.
- Grouping intervention: The central intervention holds supplied keys and ordered 160-word evidence fixed while injecting false splits or merges, aggregating within groups, and emitting one record per group.
- Controlled panel: The controlled campaign panel contains 48 fictional product pairs, with four cross-genre records tied either to one campaign root or to four independent roots.
5 RESULTS
The results separate complementary-content retention from partition effects: grouping preserves added content, false splits increase influence, and false merges reduce it under content-fixed controls. Broader interventions show checkpoint-dependent signs and substantial order sensitivity.
- Central causal decomposition: Table 1 decomposes evidence weighting into complementary-content retention, false-split inflation, and false-merge suppression while holding supplied operational keys fixed.
- Content retention: 2.91–29.82 points: adding complementary content within one group raises attack-side probability relative to the one-window endpoint.
- Partition effects: 10.27–32.66 points: falsely splitting one supplied unit raises attack-side probability, while 9.13–31.79 points: falsely merging four units suppresses influence.
- Matched control: The matched six-slot control fixes repeated non-content fields and tokenizer length, isolating content-bearing record placement under equal serialization counts.
- Multiplicity sensitivity: Bonferroni-corrected intervals retain positive lower bounds for all 16 content-fixed split/merge effects and 11 of 16 matched six-slot effects.
- Controlled campaign panel: The same-root split effect is positive for Qwen3-8B, Qwen3-4B, and Phi-4-mini but negative for Mistral-7B, establishing checkpoint-dependent behavioral signs.
- Order sensitivity: Independent-root merge loss is positive for Qwen3-4B, Phi-4-mini, and Mistral-7B but reverses for Qwen3-8B; its order interaction ranges from +41.43 to −49.92 points.
- Decision shifts: Across six checkpoints, four copies of one supplied unit shift attack-side decisions by 9.24–51.98 points, while raw copying moves leading answers by 15.62–20.62 points across three checkpoints.
6 DISCUSSION
The discussion frames partitioning as an evidence-accounting decision before generation. The proposed representation preserves group–content state, aggregates complementary content, and normalizes group-level influence while leaving checkpoint-specific behavior empirical.
- Evidence accounting: Partitioning chooses an evidence-weighting rule whenever a system chunks, duplicates, groups, or serializes retrieval results.
- Pipeline design: A source-aware pipeline separates dependence partition discovery, record grouping, complementary-content aggregation, and group-level weight normalization.
- Representation guarantee: An invariant summary preserves the full group–content state exactly when its recoding is injective, while scalar group counts discard distinctions between evidence states.
7 CONCLUSION
The paper formalizes a supplied partition as a pre-generation representation that preserves group–content distinctions while removing within-group canonical copies. Its theorems and propositions characterize invariance, faithful recoding, and limits on recovering or transporting downstream effects.
- Conclusion: The supplied partition changes model-facing influence by preserving group–content distinctions while removing within-group canonical copies.The conclusion connects the representation to exact copies, complementary passages, content-fixed split/merge interventions, and controlled campaign renderings.
- Conclusion: The two-stage identity permits dependent stages and direction-asymmetric errors, while prompt transport can still have worst-case absolute error at least L.The lower bound holds when auxiliary distributions are identical but end-to-end effects are opposite.
- Conclusion: Deterministic text-only provenance recovery exists exactly when observationally identical worlds have the same oracle provenance partition.Thus visible text alone cannot distinguish worlds whose oracle partitions differ but whose complete text observations coincide.
- Conclusion: Theorem 1 characterizes invariant summaries through a quotient that removes record order, opaque-label renaming, and within-group canonical duplicates.The quotient preserves complementary content and keeps nontrivial splits and merges distinct.
- Conclusion: A universally sufficient enforcing representation is equivalent to an injective recoding of the quotient state.This provides the paper’s formal criterion for preserving all relation-invariant information.
B.3 FAITHFULNESS, MINIMAL SUFFICIENCY, AND ADDITIVE REALIZATIONS
This section establishes that quotient-faithful representations are precisely the minimally sufficient recodings of invariant group–content states. It also gives an additive realization criterion based on finite integer-linear independence and an enforcement condition for model-visible renderers.
- Faithfulness and minimal sufficiency: The quotient representation is universally sufficient because every relation-invariant observable factors through it.Any universally sufficient representation must itself determine the quotient state.
- Faithfulness and minimal sufficiency: For an enforcing representation, universal sufficiency, quotient faithfulness, and injectivity of its quotient recoding are equivalent.The equivalence identifies the exact information-preservation requirement.
- Additive realizations: An additive realization is quotient-faithful if and only if its group-aggregator vectors are finitely integer-linearly independent.Collisions in additive outputs correspond exactly to nonzero integer relations among the aggregator vectors.
- Additive realizations: Canonical coordinate vectors provide a faithful additive realization by exposing every finite multiset coefficient.The construction uses X = R(Cont) and a(A) = e_A.
- Prompt-visible enforcement: Uniform invariance of every binary downstream map holds exactly when the model-visible renderer factors through the operational quotient.If equivalent configurations render differently, a deterministic binary rule can distinguish them.
- Prompt-visible enforcement: The formal invariance result is uniform over binary downstream maps, whereas fixed-model response is measured empirically.The default quotient preserves content and group multiplicity; authenticated identity or credibility can be added explicitly.
B.4 CONCRETE ADDITIVE CONSTRUCTION AND GROUPING-ERROR BOUND
The concrete construction aggregates canonical content within supplied groups and bounds the effect of changing partitions. The resulting content-aware discrepancy can be tighter than record-level disagreement and yields downstream decision-stability conditions.
- Concrete construction: Each group contributes once after canonical content aggregation, so adding an exact canonical copy within an existing group leaves the evidence state unchanged.The fixed canonicalizer and aggregator receive the same set for the changed group.
- Grouping-error bound: Partition-error magnitude is bounded by B times the overlap-graph discrepancy H(P, Q) when every group aggregate has norm at most B.Identical record blocks cancel, and erroneous components are summed using the triangle inequality.
- Decision stability: The sharp content-aware bound gives |m(EI(P)) −m(EI(Q))| ≤ LB∆q(P, Q) for an L-Lipschitz downstream margin.The bound is tight over the allowed scalar-aggregator class.
- Grouping-error bound: The content-aware discrepancy ∆q can be strictly tighter than H(P, Q) because changes that preserve quotient states cancel after canonicalization.In the supplied example, H(P, Q) = 4 while ∆q(P, Q) = 0.
- Decision stability: A binary decision sign remains stable whenever |m(EI(P))| > LBH(P, Q).The margin must exceed the largest admissible perturbation from the partition change.
- Natural-text evaluation: The natural-text evaluation uses externally supplied operational units, including HUMAN dependence components and WEB canonical-URL keys, with authenticated identity represented separately.The domains are analyzed separately under frozen grouping criteria.
C.5 EXTENDED CHECKPOINT AND GROUPING SPECIFICATIONS
The extended specifications test partition effects across checkpoints, controlled panels, grouping methods, and balanced order decompositions. They match serialization where needed and quantify checkpoint-dependent decision shifts and order interactions.
- Controlled generation: The controlled generation panel crosses BASE, raw four-copy, and prompt-rule conditions with two attack sides and two record orders across frozen HUMAN and WEB items.All three formal runs produced 960/960 valid predictions per checkpoint.
- Partition intervention: The primary grouping audit fixes supplied keys, holds evidence words and order constant, and changes only emitted partition and serialization.Its outcome is attack-side likelihood share, with strict selection, half-tie selection, and likelihood margin as secondary endpoints.
- Matched serialization: The Phase 11 control places the same ordered 160-word stream across one or four nonempty slots within an identical six-slot skeleton.Pair padding equalizes final characters, bytes, and tokenizer lengths while preserving model-visible empty slots and headers.
- Controlled campaign panel: The 48-item campaign panel compares four same-root genre-varied records with four independent-root records, using exactly matched 160-word streams.Each checkpoint completed 768/768 choices, and the balanced design averages attack sides and order mirrors.
- Order decomposition: For Qwen3-8B, the independent-root contrast was +41.43 points in original order and −49.92 in reverse order.The main table therefore reports balanced mirrored effects while retaining both reversals whose pointwise intervals exclude zero.
- Grouping specifications: Exact partition recovery succeeds for oracle identifiers and embedding clustering, whereas exact hashing and MinHash split every controlled same-root set.On 96 same-root scopes, the lexical methods have pairwise recall and F1 equal to zero, with 576 false-split pairs.
D.3 GROUPING-KEY ROBUSTNESS CURVES
The grouping-key robustness analysis varies emitted group count and fixed per-group content budget, while multiplicity-adjusted results assess the stability of the central contrasts. The fixed-content split/merge effects retain positive adjusted support across all cells.
- Robustness curves: The robustness curves jointly vary grouping and fixed per-group content budget across partial-copy and operational-unit panels.Content-fixed contrasts in Table 1 provide the primary comparison.
- Multiplicity sensitivity: 16/16 content-fixed split/merge cells retain positive lower bounds under ordinary and familywise 95% intervals.Across all 24 Table 1 cells, 23 retain positive lower bounds after multiplicity adjustment.
- Multiplicity sensitivity: 11/16 matched six-slot control cells retain positive lower bounds after familywise adjustment.The adjusted full-content exception is the Phi-4-mini/WEB gain of +2.91 points with interval [−0.02, +5.95].
D.7 PHASE 7 PRIMARY CONTRASTS
The Phase 7 contrasts evaluate same-root repetition and presentation effects using strict selection, likelihood, provenance, and generation diagnostics. Repetition sensitivity is consistent in direction, while prompt recovery and selectivity vary by checkpoint and domain.
- Primary contrasts: Same-root repetition contrasts are evaluated as equal-domain item means in percentage points with 95% bootstrap intervals.A pass requires an estimate of at least 5 points and an interval lower bound above zero.
- Checkpoint effects: Phi’s attack remains positive in both domains, all three presentations, and both record orders.Its six presentation-by-order cells span 4.35 to 48.64 points.
- Primary contrasts: Across eight model-domain cells, replication invariance is 0/8 and Four-Unit–Base sensitivity is 8/8.Four-Unit–Copy selectivity is 4/8, while prompt recovery is 3/8.
- Threshold sensitivity: Copy–Base lies outside equivalence in all eight cells across every tested margin.Four-Unit–Base remains positive in eight of eight cells through the 7.5-point margin and seven of eight at 10 points.
- Generation diagnostics: All generated trials have valid leading answers, so all-trial and valid-output estimates are identical.Agreement measures consistency between candidate and generated interfaces, whose Copy–Base effects share direction.
E RUN LINEAGE AND REPRODUCIBILITY
The run lineage freezes inference settings, trial accounting, and bootstrap procedures, then independently rechecks reconstruction, scoring, aggregation, and reported statistics. The recomputation artifacts pass the stated verification checks with zero failures.
- Frozen configurations: Inference precision is fixed per checkpoint, with deterministic candidate scoring, greedy generation, native termination sequences, and frozen bootstrap seeds.Model-specific chat templates remain fixed during compilation.
- Trial accounting: The run lineage records complete trial counts across original phases, added checkpoints, controlled generation, and grouping stress tests.Phase 8 and the extended candidate grid each report 3,824/3,824 completed trials per applicable configuration.
- Independent verification: Independent scripts recheck prompt reconstruction, likelihood arithmetic or answer parsing, labels, aggregation, dependence components, and bootstrap statistics.The grouping, added-checkpoint, generation, serialization, and controlled-GEO audits pass with zero errors.
- Reproducibility artifacts: The recomputation artifacts pass fresh-extraction verification, including all 3,072 Phase 12 predictions and an independently recomputed maximum numerical difference of 5.33×10−15.The manifest covers 284 non-manifest files in a 285-file archive.