Source-linked AI summary
Portable Semantics, Private Dialects: Reuse and Negative Transfer in Latent Communication Between Language-Model Cells
Narcis Marincat
TL;DR
The paper asks whether independently trained latent communication interfaces share a reusable packet language and whether inherited interfaces help later learning. It uses sealed causal interoperability, failure-localization, and adaptation experiments, finding initialization-stratified dialects, operator-side zero-shot failure, and severe negative transfer from one globally trained interface. These conclusions are bounded to a near-transfer 17-state setting and checkpoint-specific adaptation tests.
Problem
The study asks whether independently trained societies share one packet language, where strict zero-shot transfer fails, and whether inherited interfaces accelerate, leave unchanged, or harm learning.
Method
A preregistered four-part campaign audits all directed society pairs with sealed held-out alignment tests, localizes zero-shot failure with a source-span control, reinitializes interface components factorially, and replicates inheritance tests on second target streams.
Results
The six interfaces do not form one raw language: one same-initialization pair is exactly interoperable, another is partially compatible, and all 26 cross-initialization directions fail; reinitialization raises accuracy from 0.169 to 0.857.
Takeaways & Limitations
Restricted training yields semantically portable interfaces expressed in initialization-anchored coordinate dialects, whereas the tested globally trained interface can become a severe negative-transfer prior.
Takeaways & Limitations
Conclusions are bounded to a near-transfer 17-state setting; the negative-transfer factorial uses one globally visible checkpoint, and P2 uses one sealed start-value formulation.
Abstract
from arXiv · showhide
In shared-genome language-model societies, restricted evidence visibility favors reusable, value-indexed latent packet interfaces, whereas the sole high-performing globally visible model in the parent study learned an episode-entangled code. This companion study asks whether independently trained societies share one packet language, where strict zero-shot transfer fails, and whether inherited interface state helps or harms later learning. First, a leakage-controlled causal interoperability audit over all 30 ordered pairs of six independently trained restricted societies -- under sealed held-out structure and a preregistered raw/orthogonal/linear/nonlinear alignment ladder -- shows the six semantically similar interfaces do not form one raw language: one same-initialization pair is exactly interoperable in both directions, a second shows asymmetric partial compatibility, and all 26 cross-initialization directions fail every frozen alignment rung. Second, within the tested decomposition and a single sealed source formulation, a source-span control localizes strict zero-shot failure to interpretation and execution of the new operator instructions. Third, in a matched adaptation factorial, the globally trained communication interface acts as a severe negative-transfer prior: reinitializing only the packet reader, writer, and mouth raises final depth-three accuracy from 0.169 to 0.857. Fourth, across two restricted checkpoints and two independently frozen target streams each, inherited interfaces never exceeded fresh-interface controls by the preregistered 0.10 margin. All primary conclusions are bounded to a near-transfer 17-state setting; the negative-transfer factorial concerns one globally visible parent-cohort checkpoint, while an appendix adds a post hoc tagged-global twin case study.
1 Introduction
This study tests whether learned latent communication interfaces are interoperable, reusable, and beneficial to inherit across new tasks. A preregistered campaign finds initialization-stratified dialects, localizes zero-shot failure, and identifies severe interface-specific negative transfer.
- Research questions: The study asks whether independently trained societies share a packet language, where zero-shot failure occurs, and whether inherited interfaces help or harm adaptation.These questions concern interoperability, failure localization, and transfer value.
- Campaign overview: The study combines a sealed interoperability audit, zero-shot failure localization, component-reinitialization factorial, and second-stream inheritance replication.Experiments use released checkpoints under sealed manifests, frozen decision rules, and bit-exact machine admission.
- Interoperability: The six semantically similar interfaces do not form one raw language: one same-initialization pair is exact, another is partially compatible, and all cross-initialization directions fail.The alignment audit covers all 30 ordered pairs under raw, orthogonal, affine-linear, and nonlinear frozen mappings.
- Transfer: Reinitializing the globally trained checkpoint’s reader, writer, and mouth raises final depth-three accuracy from 0.169 to 0.857.The result is a checkpoint-specific negative-transfer case study, not a population-level estimate.
- Novelty: The contribution combines causal interoperability testing of emergent communication with a matched component-reinitialization factorial that localizes negative transfer to the learned interface.The authors distinguish this combination from prior individual ingredients.
2 Setting and prior results
The societies use four communicating language-model cells with a shared frozen genome and LoRA, exchanging continuous packets for ordered function composition over Z17. Prior results show restricted societies learn portable value-indexed interfaces, while strict whole-system transfer to a new instruction family fails.
- Society setting: Each society has four cells sharing a frozen Qwen2.5-0.5B-Instruct genome and rank-8 LoRA, with communication restricted to continuous packets.Each packet contains two 896-dimensional vectors, and the final cell decodes through a frozen LM head plus learned mouth projection.
- Task: The task composes ordered natural-language functions over Z17, with held-out operator programs and phrasings used for evaluation.Chance performance is 1/17 ≈0.0588.
- Prior results: All six restricted societies expose approximately value-indexed interfaces, with within-checkpoint same-value packet transplants preserving behavior at 0.94–1.00.Counterfactual-value packets redirect outputs toward mathematically predicted answers.
- Prior results: Strict zero-shot transfer to Family D fails at chance, although restricted societies preserve value-packet behavior and mouth decoding at 1.000.Family D uses new affine transformations, vocabulary, and phrasing distribution.
- Experimental discipline: The experiments were frozen in a written mini-plan with sealed manifests, committed splits, hashed scripts, and bit-exact machine admission.The manifest fixed checkpoint hashes, mapper specifications, seeds, and held-out evaluation structure.
3 Related work
The related work spans causal interchange, cross-model alignment, latent communication, emergent conventions, initialization effects, and negative transfer. The paper distinguishes its contribution by combining these ingredients around causal transplantation of learned communication packets.
- Alignment and causality: Prior alignment and causal-interchange work motivates testing whether internal states encode transferable high-level variables, while functional compatibility work studies cross-model mappings.The paper’s audit differs by testing emergent communication messages behaviorally in recipient computation.
- Positioning: The paper’s distinctive conjunction is communication-specific causal transplantation across all directed pairs, with known semantic values, held-out fitting and evaluation, frozen maps, and derangement controls.This design separates functional packet interoperability from geometric alignment alone.
- Latent communication: Latent communication research includes universal activation interfaces, orthogonal transformations, causal channel audits, and hidden reusable APIs.These precedents cover alignment, controls, and reusable internal interfaces without establishing the paper’s combined claim.
- Conventions and initialization: Emergent-communication studies establish that independently trained partners can adopt incompatible conventions, while initialization, data order, and permutation symmetries affect learned representations.The paper places its initialization-stratified dialect analysis within this broader convention literature.
- Negative transfer: Negative-transfer research shows that source training can impair target learning and that partial reinitialization can localize which components carry transfer.Related work also links warm starts to reduced plasticity and target-side reader adaptation.
4 P1: a sealed cross-model transplant audit
P1 evaluates whether six restricted societies can causally exchange natural packets under sealed held-out tests and a frozen alignment ladder. Exact interoperability appears only within one same-initialization pair; partial compatibility appears in another, while all cross-initialization directions fail.
- 4.1 Design: The audit harvests packets by running value and position, separates fit and test banks, and evaluates held-out values, programs, phrasings, positions, and packet instances.The alignment ladder tests raw identity, orthogonal, affine-linear, and nonlinear maps under frozen gates.
- 4.1 Design: The pure-carrier layout makes same-value and counterfactual conditions the same indexed-packet-following test under different target values.At k = 0, the source packet is deterministic per value, effectively providing one donor instance per class.
- 4.2 Results: Recipient self-diagonals are 1.000, while shuffle controls average 0.052 at the raw rung and 0.006 across fitted-map variants.No passing cell’s shuffle approaches its paired indexed-packet score.
- 4.2 Results: The initialization-204 pair is exactly raw-interoperable in both directions, scoring 1.000 in every test cell across both instruction families with 0.000 shuffle controls.Table 1 reports same-value transfer accuracy and shuffle controls for the same-initialization directions.
- 4.2 Results: The initialization-203 pair shows asymmetric partial raw compatibility, with minimum cells of 0.875 and 0.750, so neither direction passes the strict floor.The pair is reported as partial compatibility rather than a near-miss.
- 4.2 Results: All 26 cross-initialization directions fail the raw, orthogonal, affine-linear, and width-64 nonlinear alignment gates on sealed semantic classes.Position-stratum transfer can be high, but fitted maps fail semantic extrapolation to held-out values.
5 P2: localizing the zero-shot failure
Source-span controls show that, within the tested decomposition and one sealed Family-D formulation, zero-shot failure occurs during new-operator interpretation and execution rather than source parsing or packet/output handling.
- Source-span control: 1.000 source-only carrier accuracy shows all six restricted checkpoints correctly parse Family-D start values and publish the corresponding source packets.Family-A source and teacher-forced packet controls also scored 1.000.
- Method: P2 uses sealed source-only Family-D episodes, matched Family-A controls, and teacher-forced source-packet controls scored across all 17 values.
- Localization: Strict zero-shot failure is localized to interpreting and executing new operator instructions, not source-state parsing, packet carriage, or final value-to-label decoding.This conclusion combines the source-only result with ceiling packet and mouth portability and chance-level exhaustive operator execution.
- Scope: The localization is bounded to the single sealed Family-D start-value formulation and is not evidence about paraphrase generalization.
6 P3: a matched interface-adaptation factorial
A matched factorial shows severe interface-level negative transfer on one globally visible checkpoint: reinitializing only the communication interface lets the same trained LoRA learn Family D, whereas inherited initialization remains trapped.
- Design: The factorial varies only communication-interface treatment across frozen inherited, trainable inherited, and fresh-interface arms while holding the LoRA, architecture, optimizer family, and budget fixed.All arms are also evaluated with packets cut, and the primary endpoints are final held-out depth-three accuracy and normalized learning-curve AUC.
- Results: Final depth-three accuracy rises from 0.169 with the fully adaptable inherited interface to 0.857 with a fresh interface using the same globally trained LoRA.The preregistered differences are Δfinal(c−a)=+0.7927 and Δfinal(c−b)=+0.6883, with corresponding AUC gains of +0.359 and +0.365.
- Results: The fully trainable inherited interface remains trapped at 0.169 within 5,000 updates, whereas reinitialization produces a sharply different optimization trajectory consistent with basin dependence.
- Scope: This is a checkpoint-specific mechanistic case study and does not establish that every globally trained interface behaves similarly.
- Interpretation: The fresh-interface 0.857 endpoint shows retained LoRA capacity to acquire Family-D operator semantics after decoupling from the inherited interface.
7 P4: second-stream replication of the inheritance comparison
Across two restricted checkpoints and two independently frozen target streams per checkpoint, inherited interfaces again failed to achieve the preregistered large advantage over fresh interfaces.
- Design: P4 repeats the inherited-versus-fresh comparison on a second independently frozen target stream for each checkpoint, retaining the +0.10 criterion in both metrics.
- Results: The inherited interface never exceeded the fresh-interface control by the required 0.10 margin in normalized AUC or final depth-three accuracy.This result holds across both restricted checkpoints and both target streams per checkpoint.
- Interpretation: The stream-replicated outcome excludes the preregistered large advantage, not an exact null: AUC differences range from 0.001 to 0.083 and final differences from −0.042 to +0.025.
8 Discussion
The campaign separates semantic portability, raw coordinate alignment, and adaptation value. Restricted interfaces were semantically reusable but expressed in initialization-anchored dialects, while the observed global interface produced severe negative transfer; conclusions remain narrowly scoped.
- Reinitializing the globally trained reader, writer, and mouth while retaining the shared LoRA removed the observed adaptation failure, linking interface state to negative transfer.This is a checkpoint-specific mechanistic case study rather than a population-level law.
- The study supports a mechanistic link between semantic canonicality and interface reusability, but does not establish that canonical interfaces always transfer or global interfaces generally cause harm.
- All six restricted interfaces transported the Family-A value code into a new linguistic and operator context, but they did not form one raw cross-model language.The interfaces were semantically portable while retaining initialization-anchored coordinate conventions.
- The conclusions are bounded to a near-transfer 17-state setting, a single globally visible checkpoint, tested alignment families, and a single sealed start-value formulation.The inheritance comparison excludes only the preregistered large benefit, and richer alignments were not tested.
- Future experiments should vary target reinvention cost and source diversity to classify positive, neutral, and negative interface transfer prospectively.The current negative-transfer result motivates treating negative transfer as a first-class hypothesis.
Reproducibility and artifacts
The study releases its computational artifacts and records the controls needed to reproduce the analyses. Checkpoints, adapted runs, seeds, metadata, and frozen evaluation materials are identified in the companion repositories.
- The experiments used released checkpoints, pinned software and genome revisions, bit-exact machine admission, and a sealed P1 manifest established before held-out scoring.
- The twin-dialect release includes RNG seeds, donor and recipient metadata, per-case predictions, and the complete self/cross transplant matrix.The companion repository also hosts parent, tagged-global, and adapted-run checkpoints plus fitted alignment artifacts.
A Post hoc case study: twin dialects across training regimes
The post hoc twin study compares a restricted and role-tagged global society sharing initialization and episode order but differing jointly in evidence visibility and tagging. Restricted packets were value-canonical, whereas global packets were context-dependent, with partial late-interface compatibility but no evidence of a universal language.
- Setup: The twins shared byte-identical initialization and episode order, but the regimes differed jointly in evidence visibility and ownership tagging.The restricted twin scored 0.8775/0.6713 at depths two/three, while the tagged-global twin scored 0.1593/0.7914.
- Method: The transplant audit harvested natural packets from correctly answered depth-three episodes across all 17 values and two phrasing rotations, then substituted packets at each interface.The same 120 success-conditioned recipient episodes were used across donor models, with intact accuracy 1.000 by construction.
- Findings: The restricted twin’s packets were causally sufficient for the running value across all three interfaces under natural, counterfactual, and deletion interventions.Same-value substitutions preserved behavior, counterfactual values redirected predictions, and deletion or mismatched values reduced performance toward chance.
- Findings: The final interface retained partial bidirectional interoperability, whereas the twins’ early interfaces were raw-incompatible.
- Findings: Recipient-normalized late-interface compatibility was nearly symmetric at 0.650 versus approximately 0.647, consistent with predominantly recipient-side limitations.The ratio analysis is not a formal decomposition of donor and recipient effects.
- Interpretation and scope: The case does not establish a shared packet language or isolate initialization as the cause of late compatibility, because the twins share several other pressures.The study also does not measure geometric alignment between their packet spaces.
- Novelty boundary: The work claims a novel combination of same-lineage cross-regime training, raw message transplantation, semantic counterfactual scoring, and recipient-normalized self-interchange analysis.