Source-linked AI summary
Scaling Model-Generated Distillation Data Can Make Latent Teacher Traits More Recoverable
Zhichen Dong, Zhixuan Liu, Yuyu Fan, Xiangtian Li, Shuyang Zhang, Chao Yang
TL;DR
The paper asks whether scaling model-generated distillation data can reveal latent teacher-specific traits that ordinary filtering may miss. Using induced teachers, filtered off-task datasets, matched controls, and separate-domain behavioral readouts, it finds that larger independent datasets generally make the intended trait more detectable and specific. The authors therefore recommend trait-aware curation and scale-matched evaluation, while noting that the controlled setup and readouts limit generalization.
Problem
Existing distillation views treat more generated data mainly as a performance lever, while evidence is limited on how scale affects hidden teacher-specific signals in students.
Method
The study induces teacher traits, generates filtered off-task data at multiple independent sample sizes, trains students with matched no-trait controls, and evaluates separate-domain behavioral readouts.
Results
Larger independent datasets generally make induced traits more recoverable and target-specific, including under LoRA analysis, competing traits, richer perturbations, and cross-model transfer.
Takeaways & Limitations
Generated-data scaling should be paired with trait-aware curation, matched controls, and evaluations conducted at deployment scale.
Takeaways & Limitations
The experiments are controlled probes rather than full industrial distillation reproductions, and the chosen behavioral probes may miss signals outside their prompts or scoring rules.
Abstract
from arXiv · showhide
Scaling model-generated data is usually viewed as improving distillation: more examples should increase coverage, reduce noise, and produce stronger students. We show a second effect: larger datasets can make subtle teacher-specific signals easier to detect in the trained student, even when examples are off-task and never mention the trait. In a controlled setup inspired by subliminal learning, a teacher induced to express a target trait generates restricted off-task data, such as number-only completions. Students trained on different amounts of independent off-task data are evaluated in a separate domain, with matched no-trait controls isolating target-specific transfer. Our main finding is that larger independent datasets make the teacher's induced trait stand out more clearly in the student's later behavior. Other plausible traits may also strengthen with scale, but the target usually grows more. When the small-scale student already favors the target, scaling mainly amplifies that behavior; when it favors a related or salient alternative, more data can shift behavior toward the intended trait. Analyses of learned LoRA updates show a parallel trend. These effects appear across model families, trait types, multi-trait settings, and cross-model transfer. Our results suggest that scaling generated distillation data should be paired with trait-aware curation and evaluation, even when the data appears off-task or benign.
1 Introduction
The paper identifies a second effect of scaling model-generated distillation data: independent off-task examples can make latent teacher traits more recoverable in students. Controlled transfer experiments show that scaling often sharpens coarse or ambiguous behavior toward the intended trait, including in multi-trait and cross-model settings.
- Model-generated distillation data can carry subtle teacher-specific signals beyond its surface task, even when examples contain no explicit trait cues.
- The study induces a target trait in a teacher, generates filtered off-task examples, trains students at different scales, and uses matched no-trait controls.Students are evaluated in a separate domain to isolate transfer from data format and training effects.
- Larger independent datasets usually make the intended trait grow more than plausible alternatives, sharpening transfer from coarse attractors such as lion-like or salient plant preferences.At small scales, behavior may favor related alternatives; scaling can instead resolve it toward the target.
- The pattern extends beyond isolated single-trait transfer: competing traits can become more visible with scale, while cross-model transfer remains noisier after background differences are controlled.
- Scaling generated data should be paired with trait-aware curation, matched controls, and evaluation at deployment scale because small pilots may underestimate learnability.
2 Related Work
Prior work establishes model-generated supervision, hidden information in model outputs, and subliminal behavioral transfer. This paper studies how increasing the number of independent generated examples changes which latent teacher signals become learnable.
- Knowledge distillation has expanded from labels, logits, and softened predictions to generated instructions, explanations, rationales, preferences, and agentic trajectories.
- This paper adds the question of whether more independent generated examples change which latent teacher-specific signals become learnable by the student.
- Model outputs can leak information beyond intended semantics through privacy attacks, memorization, soft-label distillation, and related channels.
- Subliminal learning shows that filtered, semantically unrelated teacher outputs can lead students to express the teacher’s induced trait later.
3 Distillation Setup and Trait Evaluation
The paper measures trait transfer by training matched students on filtered off-task datasets from induced and no-trait teachers, then evaluating behavioral readouts in a separate domain. Reference adjustment and, for preference traits, closed-set localization quantify target-specific recoverability across data scales.
- Distillation protocol: A trait-bearing teacher and matched no-trait teacher generate independently sampled off-task datasets using identical prompts, restrictions, and target-excluding filters.The filter can enforce carriers such as number-only completions while removing explicit mentions of the target trait.
- Distillation protocol: Students use the same base model and training algorithm, while n counts independently generated accepted examples rather than repeated exposure to a fixed dataset.
- Trait evaluation: Behavioral readouts measure each target in a separate domain, using domain-specific preference scores, GSM8K accuracy, or posture scores.
- Trait evaluation: The reference-adjusted target score measures target behavior above a matched no-trait student after accounting for background effects of format, filtering, model family, and training recipe.
- Trait evaluation: For animal and plant preferences, the localization margin Γτ(n) compares the target’s adjusted score with the strongest non-target candidate in a fixed closed set.A positive margin means the target exceeds every non-target candidate; this metric is not used for task- or posture-style traits.
- Interpretation: Recoverability denotes increased detectability under a prespecified behavioral readout, not recovery of the teacher’s full internal state or every hidden property.
4 Scaling Off-Task Distillation Data Increases Readout Recoverability
The paper tests whether increasing independent off-task carrier data makes induced teacher traits more visible and specifically localized in students. Across preference, model-family, update-space, task, and safety settings, scaling generally strengthens target readouts and can shift behavior away from salient alternatives.
- 4 Scaling Off-Task Distillation Data Increases Readout Recoverability: The study varies independent off-task carrier examples, fixed closed-set readouts, LoRA update probes, teacher perturbations, and model families to measure recoverability.The protocol separates independent-data scale from repeated exposure and examines behavioral and update-space correlates.
- 4.1 Single-Trait Preference Scaling: 2/16 to 14/16 animal targets and 2/16 to 11/16 plant targets achieve positive localization margins as independent data increases.The ranges are reported over 1K–40K animal examples and 10K–100K plant examples, respectively.
- 4.1 Single-Trait Preference Scaling: Target readouts can rise before the intended candidate is top-ranked, so amplification and localization measure distinct aspects of recoverability.Plant targets may improve while shared attractors remain stronger at intermediate scales.
- 4.1 Single-Trait Preference Scaling: With total training rows fixed at 60K, increasing unique carrier examples improves readouts, arguing that independent sample count matters beyond repeated exposure.This control holds optimization exposure constant while varying carrier diversity.
- 4.2 Update-Space Landscape Analysis: Higher-scale target adapters move toward target-favoring update-space regions and larger target–attractor margins, indicating changed update direction as well as stronger adapter scaling.Some low-scale rays cannot be rescued merely by increasing adapter scale.
- 4.3 Scope Across Teacher Perturbations: Scale-dependent transfer extends to GSM8K/MATH readouts and safety posture, with unsafe rates rising from 2.0% to 33.7% for LLM judges and 3.3% to 38.0% for human judges.Format-only and shuffled controls are flatter, while refusal-marker probability remains essentially unchanged in the safety evaluation.
5 Beyond Isolated Single-Trait Transfer: Interactions and Cross-Model Transfer
When multiple traits share a carrier, they compete and their recovery depends on both data scale and carrier composition. Cross-model transfer remains detectable after matched controls, but is noisier and trait-dependent.
- 5.1 Multiple Simultaneous Latent Traits: Multiple traits compete under shared carrier data, so prompt order and carrier composition shape which traits emerge.The first-mentioned trait resembles single-trait transfer, while the later trait is weaker at smaller scales.
- 5.1 Multiple Simultaneous Latent Traits: Larger carrier datasets gradually reveal traits initially suppressed by competition, while mixing prompt orders produces more balanced joint recovery.Figure 6 compares animal-first, plant-first, and mixed carrier data; node labels count pair-seeds where both targets become top-ranked.
- 5.2 Cross-Model Transfer and Reference Controls: Cross-model transfer remains scale-dependent, but model mismatch introduces background shifts and makes animal and plant effects weaker and noisier.Matched no-trait cross-model references are used because unperturbed source data can already shift student behavior.
- 5.2 Cross-Model Transfer and Reference Controls: All five induced identities become top-ranked after cross-model distillation, including GPT for a 100K Gemma student trained from a Qwen teacher induced toward GPT.Similar scale-dependent trends are reported for Qwen2.5-7B-Instruct→Qwen2.5-3B-Instruct and Llama→Qwen.
- 5.2 Cross-Model Transfer and Reference Controls: The study does not identify which factors beyond data scale determine cross-model transfer strength.Model mismatch may attenuate transfer but is not treated as a reliable safeguard against unwanted teacher-induced behavior.
6 Conclusion
The paper shows that scaling model-generated distillation data can make hidden teacher traits more recoverable, even when student-visible examples are off-task and exclude the target. This pattern extends across richer trait settings and model families, motivating trait-aware evaluation at deployment scale.
- 6 Conclusion: Scaling distillation data can make latent teacher traits more recoverable from off-task, target-excluding examples.Larger independent carrier datasets amplify target readouts and can sharpen coarse attractors toward intended traits.
- 6 Conclusion: The pattern persists with more noise under competing dual-trait preferences, richer teacher perturbations, and cross-model transfer.Analyses of learned update space show a parallel increase in target-related structure.
- 6 Conclusion: Generated-data scaling should be paired with trait-aware curation, matched controls, and evaluations at deployment scale.
7 Limitations
The experiments provide controlled evidence about scale-dependent hidden transfer, but their settings, metrics, mechanisms, and mitigations do not fully represent industrial distillation.
- 7 Limitations: The experiments are controlled probes rather than full reproductions of industrial distillation pipelines.They use restricted carriers, explicit trait induction, LoRA SFT, matched references, and targeted evaluation domains.
- 7 Limitations: The closed-set localization metric applies only to preference traits with fixed candidate sets, while behavioral probes may miss signals outside chosen prompts or scoring rules.
- 7 Limitations: Update-space analyses correlate geometry with recoverability but do not explain how off-task outputs encode teacher traits, and the paper proposes no complete mitigation.
- 7 Limitations: The findings motivate broader tests in realistic data mixtures and training pipelines.
8 Ethical Considerations
The paper treats hidden behavioral transfer from harmful teacher conditions as a dual-use risk requiring controlled experimentation and careful artifact handling. It recommends separating, redacting, and auditing safety-related materials and checkpoints.
- 8 Ethical Considerations: Safety experiments use a teacher-only misalignment prompt and fixed probes to test residual student-side behavioral differences after target-excluding training.The prompt and probe details are included to make this methodological claim auditable.
- 8 Ethical Considerations: Student-training rows exclude the misalignment prompt and contain filtered number-continuation prompts with accepted number-only completions.Open-ended probes use short-answer caps and aggregate answer/refusal or unsafe labels.
- 8 Ethical Considerations: Harmful-teacher examples should be redacted or excluded from public releases, and derived checkpoints should not be deployed.Follow-up work should report separation among hidden prompts, carrier data, and evaluation artifacts.
- 8 Ethical Considerations: Apparently benign generated data may transmit unsafe preferences, backdoorlike behaviors, or demographic, cultural, and multilingual biases.Pipelines should track provenance and source-model conditions and audit traits at intended usage scale.
A Experimental Resources and Setup Details
The experimental resources include protocol-generated carrier and readout sets, standard evaluation datasets, and multiple adaptation-regime ablations. These experiments test whether scale-dependent transfer persists beyond the default LoRA setting.
- Experimental resources: Carrier banks and animal/plant readout sets are generated or specified by the protocol rather than imported public datasets.
- Evaluation resources: GSM8K supplies task data for math-teacher construction and downstream math readouts, while MATH provides harder held-out mathematical evaluations.
- Safety resources: XSTest, HEx-PHI, and BeaverTails support safety-posture, unsafe-prompt, and harmful-response diagnostics in the safety experiments.
- Adaptation ablations: DoRA, OPD, multiple LoRA ranks, and full-parameter SFT test whether scale-dependent transfer depends on adaptation method or capacity.
- Adaptation ablations: By 40K carriers, DoRA makes each of four animal traits top-ranked, while full-parameter SFT transfer strengthens with scale when using richer AlpacaEval carriers.
B.2 Ablation: Generalization Across Carrier Distributions
Transfer generalizes beyond deterministic number-only sequences to mathematical reasoning traces and general instruction-following data. Across these practical carrier distributions, target preference transfer increases with scale without materially changing held-out math accuracy.
- Carrier distributions: GSM8K reasoning traces and Alpaca instruction–response data extend preference-transfer tests beyond number-only carriers, with target terms excluded from student-visible rows.
- GSM8K reasoning carrier: Mean Animal ∆gap remains between +2.45 and +2.58 across GSM8K carrier scales, with modestly increasing rank improvement.
- GSM8K reasoning carrier: Held-out math accuracy stays within one percentage point of the matched control at each GSM8K scale.
- Alpaca instruction carrier: From 1K to 4K Alpaca rows, mean ∆gap rises from +2.38 to +3.27 and mean rank improvement from +6.89 to +9.43.
- Generalization: Together, the results show transfer through reasoning and instruction data representative of practical distillation and fine-tuning pipelines.
B.3 Ablation: Divergence-Token Masking
Masking the first teacher–base divergence weakens transfer but leaves a strong scaling trend, showing that transferable information is distributed beyond that token. Fixed-compute analyses instead implicate the optimization trajectory’s destination as the dominant scale-dependent factor, including under cross-model transfer.
- Divergence-token masking: Transferable information remains elsewhere in carrier completions because masking weakens but does not eliminate scale-dependent transfer.
- Optimization landscape: At fixed compute, the hidden-trait gap rises from 2.033 at 10K carriers to 7.409 at 60K, while visible carrier-task improvement changes only modestly.The best nearby hidden-trait value rises from 3.275 to 10.108.
- Cross-model transfer: After matched reference adjustment, several cross-model target readouts still increase with carrier-data scale, though effects are weaker and noisier.
- Optimization landscape: A shared trait-sensitive direction improves held-out hidden-trait readouts in 11 of 12 scale–trait conditions, but its available gain changes only from 1.038 to 1.103.
- Optimization landscape: Scaling primarily changes the region reached by optimization, with hidden-trait gap at the trajectory endpoint increasing from 1.996 to 6.966.
C Detailed Experimental Setup and Scoring
The protocol generates filtered off-task datasets from trait-induced and matched no-trait teachers, trains students under controlled settings, and measures trait readouts after reference adjustment.
- Teacher and dataset construction: Trait-bearing teachers are created from base models, alongside matched no-trait teachers using the same generation protocol.
- Teacher and dataset construction: The off-task pipeline samples independently generated examples, enforces restricted output formats, and removes explicit target-trait mentions until the requested scale is reached.Scale n counts accepted examples rather than repeated exposure to a fixed dataset.
- Carrier generation: Number-carrier experiments use deterministic number-sequence completions, strict formatting, and exclusion filters before student training.
- Student training: Target and reference datasets share prompts, output restrictions, filtering rules, and accepted-example counts, while students use the same LoRA training recipe and settings.Students receive hard teacher completions with completion-only supervision, and reported conditions generally average three seeds.
- Scoring and adjustment: Students are evaluated in separate domains using trait-specific scores, including preference scores, GSM8K accuracy, or behavioral safety measures.The paper calls these behavioral evaluation scores readouts and keeps the evaluation domain separate from carrier generation.
- Scoring and adjustment: Reference-adjusted scores compare induced-teacher students with matched no-trait students to remove background effects from format, filtering, model family, and training recipe.Closed-set localization uses K = 16 animal and plant candidates; a positive margin indicates the target exceeds every non-target candidate, not open-vocabulary specificity.
D Other Experimental Protocols
Additional protocols replace or combine teacher-side perturbations while preserving the carrier-based student interface, testing task-induced behavior, safety behavior, cross-model transfer, and multiple traits.
- Task-SFT transfer: Task-SFT experiments train a teacher on GSM8K, then reuse its adapter in the same number-carrier transfer pipeline.Students do not see the original GSM8K examples or the teacher’s SFT objective.
- Task-SFT transfer: GSM8K students are evaluated with exact-match scores on held-out GSM8K and with cross-task readouts on MATH.These measurements are interpreted as readouts of task-induced behavior and answer-format residuals through the carrier channel.
- Safety transfer: Harmful and benign contexts remain teacher-local: base models, student training, and student evaluation use the default assistant system prompt.
- Safety transfer: Safety protocols generate number-carrier data from harmful or benign teacher states, then evaluate students with bounded open-question and refusal/compliance probes.The primary probe uses 300 prompts per condition, while the secondary metric compares refusal and non-refusal prefix probabilities.
- Multi-preference transfer: Multi-preference experiments perturb teachers along multiple readout axes, keep hidden teacher instructions out of student records, and test simultaneous rather than assumed independent transfer.
- Multi-preference transfer: The mixed-order condition combines disjoint carrier rows from animal-first and plant-first teachers, forming a data-level mixture rather than a new prompt or adapter merge.