Source-linked AI summary
Moving the Mean Toward the Known Good, Not Beyond It: What Inference-Time Interventions and Weight Consolidation Buy in Open-Ended Generation
Roberto I. Ono Filho
TL;DR
The paper asks whether a generation loop can improve by learning from its own verified successes rather than merely generating and selecting candidates. It evaluates inference-time interventions and LoRA consolidation on online bin packing, finding replicated gains in mean held-out candidate quality toward the known good while the best observed candidate reaches, but does not exceed, the classic heuristic. The study also reports that schematic recaps improve integration without improving development and that in-stream verifier judgments are imitated.
Problem
The paper asks what a generation loop gains from learning on its own verified successes and whether local intervention gains can compose.
Method
The study cycles through generation, verification, selection, and LoRA consolidation of verifier-approved candidates, with preregistered held-out evaluation and replicated lineages.
Results
Consolidation shifts mean candidate quality on held-out variants toward value, while the observed best reaches the classic heuristic’s level and not beyond it across three lineages.
Takeaways & Limitations
Mean quality among valid candidates can be bought and replicated, but the observed best has so far not surpassed the classic heuristic.
Takeaways & Limitations
The replication closes the original one-lineage gap for mean and value-filter effects on seven of eight registered variants, while candidate counts differ by arm and one base arm produced no valid candidate.
Abstract
from arXiv · showhide
What does a generation loop gain from learning on its own verified successes? In cycles of generate, verify, select and LoRA-consolidate on online bin packing, training on value-filtered candidates shifts what the model writes on held-out variants toward value (-1.7 points of excess, p=0.008; -3.1 against a random-consolidation control, p=0.004) while the best observed candidate converges to the classic heuristic's level and no further. A confirmation battery replicates the whole procedure three times, with fresh seeds and a never-consulted held-out set read exactly once: the mean was nearly identical in all three lineages (-2.0, -1.8, -1.9), and after aggregating within held-out variant all seven evaluable variants favored consolidation (p=0.008). The best observed candidate moved to the classic heuristic's level, exactly (0.021028 in all three lineages, for attract and for the random control alike), and never beyond it. A matched SFT-only control shows the supervised anchor, not repulsion from bad candidates, does the concentrating (96% of candidates land exactly at the classic heuristic's level). The tails cut both ways: consolidation lowers the per-candidate rate of better-than-classic candidates (10% to 3.9%) while its larger production yields more such candidates absolutely (5 against 1, on few events). As motivation we report the inference-time ledger that led here: a model-written schematic recap buys judged document integration and nothing buys development; a verifier written into the stream is imitated, 16.4 fabricated verdict lines per notebook. Mean quality among valid candidates can be bought and replicated; the observed best goes to the classic and, so far, never beyond it.
1 Introduction
The paper asks how local generation gains can compose and tests a loop adding memory, questioning, selection, verification, and consolidation. Its core contribution is evidence that LoRA consolidation on verifier-selected candidates shifts held-out generations toward known good without surpassing the classic heuristic.
- Motivation: The study asks what would make local gains in an interrupted generation loop compose.It evaluates additions that a discovering mind would have but the earlier loop lacked.
- Method: The proposed loop adds schematic memory, standing questions, candidate selection, in-stream verification, and LoRA consolidation.Preference-based repulsion and quality-diversity selection are also evaluated to preserve tails.
- Consolidation: The observed best candidate reaches the classic heuristic’s level but does not go beyond it.The result is reported alongside the held-out distribution shift.
- Replication: Three independent lineages reproduce the same mean shift, with all seven evaluable held-out variants favoring consolidation.The lineage means are −2.0, −1.8, and −1.9 points, with p = 0.008 after aggregation.
2 Related work
Related work frames the paper as self-training and preference optimization applied to verifier-selected generation, contrasted with methods that add external guidance or directed exploration. The paper emphasizes candidate-distribution tails and the trade-off between concentrating on known good solutions and preserving discovery.
- Self-training: STaR, ReST, and ReST-EM fine-tune models on self-generated outputs filtered by correctness or reward.The paper adapts this self-training pattern with LoRA and verifier-selected candidates.
- Exploration: Boundary-guided curriculum reinforcement learning and representation-based exploration improve pass@k by adding external guidance or directed exploration.These additions distinguish their reported scope from the present consolidation result.
- Preference optimization: The repulsion arms use DPO with the model’s worst candidates or known-optimum clones as rejected examples.Anchoring with supervised chosen examples sharpens the policy onto the known optimum, while unanchored training degenerates.
- Quality-diversity: The quality-diversity arm keeps one verifier-behavior elite per niche but loses value and tails in this regime.This differs from retaining only the best overall solution.
- Verified search: FunSearch and AlphaEvolve sample with an untrained proposal model and hard evaluator rather than fine-tuning the proposer.The paper reports that consolidation lowers the better-than-classic candidate rate relative to the untrained prior.
- Verifier behavior: The companion study links fabricated verifier verdicts to entrainment when verifier outputs are placed back into the prompt.This connects verifier imitation to the broader memory, tension, and judge discussion.
3 Setup
The setup evaluates interrupted generation and online bin-packing search with preregistered, cell-level judgments and paired permutation tests. It combines generated-only window assessments, whole-document readings, and verifier-based excess measurements.
- Generation loop: The generator continues each premise for 4,500 tokens with habituation enabled and interruptions every 300 tokens.The default model is Qwen3-30B-A3B-Base, while battery C uses Qwen3-8B-Base for fine-tuning.
- Judging: Generated-only 96-token windows measure surprise, connection, and coherence after injections, excluding self-copy windows.Whole streams are separately judged on integration, development, coherence, and surprise.
- Statistics: Cell means over premises are evaluated with bootstrap intervals and exact paired sign-flip permutation tests.Tests are one-sided where preregistered.
- Verifier task: Online bin packing evaluates model-written priority(item, remaining) functions by mean excess over the lower bound.Batteries S and V use near-optimal variants, while battery C uses small-item distributions where best fit is beatable.
- Document-level reading: No arm passes 2.0 on development in the document-level reading.The table summarizes whole 4,500-token streams with injected text removed and cell means over ten premises.
4 Memory and tension: the inference-time motivation
The inference-time motivation compares ways to preserve context and redirect interruptions toward unresolved content. Schematic memory improves the document-level composite over verbatim context, while the broader intervention set improves integration but not development.
- Memory: The memory arms compare verbatim context, reset reconstruction, and schema-based recaps across interrupted streams.Recaps are model-written summaries of the preceding stream, and injected text is removed before document judgment.
- Memory: +1.65 composite points favor schema over verbatim, with p = 0.001.Against reset, the registered composite contrast is not supported at +0.45, p = 0.13.
- Tension: +0.70 points favor the standing question over the neutral subject under preserved context.The question matches habituation on development, while the agenda adds nothing over schema.
- Progression: Every registered injection raises integration but never development above about 2/10.This progression wall motivates the paper’s shift from inference-time stimulation toward consolidation.
5 Selection and the verifier
Within selection, the interruption adds nothing to champion gain or candidate diversity, while stream-written verifier feedback produces little progression and frequent fabricated verdicts.
- Interruption: +0.0013 champion gain on held-out instances with interruption, versus no interruption, was not significant (p = 0.38).The reported interval was [-0.0015, +0.0047].
- Interruption: 89.5 vs 95.1 cumulative distinct valid candidates with and without interruption was not significant (p = 0.86).Under selection, the population prompt already supplies what the interruption supplied.
- Verifier feedback: +0.0003 held-out gain for verifier feedback versus the angle arm was not significant (p = 0.48).The battery used ten variants and two seeds, with baselines re-scored in text order.
- Verifier feedback: 16.4 fabricated “# Verifier:” lines per notebook appeared in the feedback arm, versus about three real ones.The authors leave their correlation with true scores and their effect on subsequent candidates open, and recommend provenance for forgeable verifier syntax.
6 Consolidation: moving the prior
LoRA consolidation on verifier-selected candidates shifts held-out candidate means toward the known classic heuristic, while the best candidate reaches but does not exceed that level. Replication and controls indicate supervised attraction drives concentration, with value filtering supplying the mean shift and production affecting absolute tail yield.
- Attraction: −0.0170 points of excess (p = 0.008) is the five-cycle attraction shift, compared with −0.0129 (p = 0.109) at the preregistered three-cycle endpoint.The cycle-2 contrast was −0.0153 (p = 0.016), and seven of eight variants improved at cycle 3.
- Attraction: 0.0311 is the trained arms’ best-candidate level, matching the classic heuristic and never exceeding it; the base fluctuates between 0.0317 and 0.0414.Figure 1 reports the same ceiling pattern across held-out variants.
- Trade-off: 10% to 3.9% is the decline in better-than-classic candidates per candidate from base to attraction, while absolute tail yield rises from 0.08 to 0.28 per notebook.The comparison rests on few events and is described as preliminary; production volume changes total yield.
- Mechanism: 96% of held-out candidates in the matched SFT-only control land exactly at the classic level, versus 100% for anchored repulsion and 72% for attraction.The control reports mean excess 0.0317, compared with 0.0311 for anchored repulsion.
7 Discussion
The interventions separate variation, mass, and integration: inference-time changes buy variation and some judged integration, while consolidation shifts probability toward verified value without moving the observed ceiling. The resulting design rules favor sampling and selection for verifier-guided search, with consolidation reserved for reliability and portfolio use.
- Variation is cheap: interruption raises fresh surprise by more than a point and multiplies valid candidates three- to fourfold.
- A schematic recap receives higher judged document integration than other arms, but the preregistered composite against reset alone is unsupported.
- Consolidation moves probability mass toward the best verifier-approved attractor within the trained family, while leaving the observed best at the classic level.
- A portfolio can combine the broad untrained prior for proposing with the consolidated prior for exploiting, but the cost-dependent fixed-budget comparison remains open.
- For verifier-guided search, the paper recommends spending on sampling and selection rather than stimulation, while keeping verifier verdicts out of prompts or measuring their fabrication.
8 Limitations
The evidence is constrained by small, specialized evaluations, limited headroom, and several design-specific implementation choices. A replication addresses the original lineage’s procedural gap for the mean and value filter, but other estimates remain narrow.
- The study uses ten premises and variants per battery, 8–30B models, a single LLM judge, and bin-packing distributions with limited headroom.
- The document judge’s absolute integration scores are low for every arm, so values such as 2.7 support relative rather than absolute claims.
- The replication uses three lineages and a fresh held-out set read once, closing the original procedural gap for the mean and value filter on seven of eight variants.
- The original primary contrast missed the registered third-cycle threshold and reached significance only after a five-cycle extension whose held-out set had already been consulted.
- Unequal candidate counts make the observed best an unequal-budget tail comparison, while far-family and tail-rate estimates rely on small samples.
- The quality-diversity arm used only three niches per variant, so its tail conclusion is specific to that implementation rather than the broader class.
9 Conclusion
Inference-time components did not consistently improve development, whereas LoRA consolidation of verifier-selected candidates shifted held-out generation toward known-good value. Across independent lineages, the mean replicated this shift, but the best candidate reached the classic heuristic’s level and never exceeded it.
- Conclusion: LoRA consolidation was the only intervention that shifted held-out generations toward value, while the combined anchored objective moved them to the classic heuristic’s level and no further.Other tested components showed no detectable development benefit on the near-saturated variants, and in-stream verification was imitated rather than providing a clear gain.
- Conclusion: 96% of candidates reached the classic level under the SFT-only control, indicating that the supervised anchor did nearly all of the concentrating.The result does not require a repulsion term from bad candidates.
- Conclusion: −2.0, −1.8 and −1.9 points of mean excess were observed across three independent lineages, with seven evaluable variants favoring consolidation at p = 0.008.Each lineage used fresh seeds and a second held-out set read exactly once.
- Conclusion: 0.021028 was the best-candidate value in every lineage for both attraction and random control, matching the classic heuristic and producing no beyond-classic candidate.The best statistic improved from the base’s 0.0243–0.0247 to the classic level.
A Reproducibility and experimental provenance
The paper distinguishes a prospective sequential extension from independent confirmation and documents the code, data, judgments, registrations, and analysis materials used to establish provenance.
- A Reproducibility and experimental provenance: The cycle-5 result is a prospective sequential extension, not an independent confirmation.The independent lineages with never-consulted held-outs constitute the confirmation battery.
- A Reproducibility and experimental provenance: Code, run data, judgments, pre-registrations and the dated laboratory notebook are available in the companion study’s repository.The repository includes scripts for the evolution, problem-loop, consolidation, DPO-LoRA, and analysis pipelines.