Source-linked AI summary

Off-Target Effects of Response-Style Alignment in a Korean 27B Language Model

Hyojung Han

arXiv:2609.11291v1cs.AI

TL;DR

The paper asks whether Korean response-style alignment changes untargeted abstention and securities disclosure, and whether those changes reflect content or emission policy. It post-trains Qwen3.8-27B with matched target-form controls and evaluates answer propensity, response length, conditional composition, and detector agreement. The changes are dominated by how often and how much the model emits, while treatment-dependent answered subsets prevent interpreting conditional contrasts as latent preference changes.

  • Problem

    The paper examines whether post-training for Korean response style changes untargeted abstention and securities disclosure, behaviours not named in the objective.

  • Method

    The study post-trains Qwen3.8-27B with matched controls that hold prompts, recipe, data volume, and serving fixed while changing target text.

  • Results

    Across seventeen contrasts, answer propensity moved by up to 13.4 pp while the composition term stayed between +0.00 and +0.45 pp; disclosure loss was largely reproduced by assigned response length.

  • Takeaways & Limitations

    Deployment-visible off-target changes in this pipeline primarily reflect how often the model answers and how much it says, not an identified change in latent conditional preference.

  • Takeaways & Limitations

    The study does not identify which individual style feature causes the bias-axis movement, and training arms contain at most three seeds.

Abstract

from arXiv · show

We post-train Qwen3.8-27B for Korean response style -- verbosity, list and markdown usage, discourse structure and register -- and measure two behaviours the objective never targets: abstention on ambiguous social questions in KoBBQ, where the benchmark-correct answer is UNKNOWN, and unprompted disclosure in securities guidance. Both move, and the changes are expressed primarily through the model's emission policy: how often it answers and how much it says. Matched target-form controls show that answer propensity depends on the training target, not the prompt set or recipe alone. Holding prompts, recipe, data volume and serving fixed and changing only the target text, three style seeds give positive answer-rate point estimates (mean +0.82 pp) and three neutral seeds negative ones (mean -1.53 pp); the observed seed ranges do not overlap and the means differ by 2.34 pp. A length-matched arm lies between them, and a fourth arm that stays short while preserving hedging is unstable across seeds, so which feature of the form is responsible is unresolved. For absolute stereotyped exposure the decomposition into an answer-propensity term and a conditional-composition term is an algebraic identity, not a finding; its empirical content is where the movement went. Across the trained checkpoints the changes are dominated by answer propensity while the composition term stays small, and because that term is evaluated on treatment-dependent answered subsets we do not read it as evidence about latent preference. Two measurement results follow. A between-arm contrast in conditional stereotyped share does not identify a change in conditional content preference when answer status is treatment-dependent. And agreement between two rule detectors for the same construct runs from 0.44 to 0.99 depending on which checkpoint produced the text -- observable without any reference labels.

1 Introduction

The paper studies whether Korean response-style alignment changes untargeted abstention and disclosure behaviours. It argues that these changes primarily reflect emission policy—how often the model answers and how much it says—rather than demonstrated changes in capability or latent preference.

  • 97.3% of base responses contain bullets, 85.1% contain markdown headings, and median length is at least 1,329 characters.These observable properties motivate post-training intended to change response form rather than safety, capability, or disclosure.
  • Response style is defined by deterministic output-form features covering verbosity, formatting, discourse structure, and register.The construct is operationalized through observable features rather than a learned composite or a human-likeness claim.
  • The study measures abstention on ambiguous KoBBQ questions and disclosure language in Korean securities guidance.KoBBQ treats UNKNOWN as correct for ambiguous contexts; the disclosure instrument tracks required warnings such as principal loss and lack of deposit insurance.
  • Style alignment moved both untargeted behaviours, with observed changes carried primarily by answer propensity and response length.The paper distinguishes emission claims from capability claims and attributes most absolute-exposure movement to emission policy.
  • Matched controls ask whether target text, rather than prompts or the fine-tuning recipe alone, accounts for the behavioural movement.The results are organized around four questions concerning target effects, emission policy, measurement validity, and composition of safety repairs.

2 Related Work

Related work frames the paper as a study of off-target regression, emission and content decomposition, verbosity confounding, and measurement validity. It also situates the experiments within Korean evaluation resources and sequential alignment methods.

  • Prior work shows benign or later fine-tuning can erode earlier safety objectives and disrupt evaluation consistency.The paper positions style post-training as a benign fine-tune whose untargeted effects are measured explicitly.
  • BBQ’s ambiguous score already factors into answer propensity and the distribution of answers among answered items.This decomposition motivates separating threshold-like answering behaviour from conditional composition.
  • The work connects to selective prediction, self-knowledge, bias-score validity, weight merging, and order-dependent sequential fine-tuning.Its Korean setting uses KoBBQ, KoSBi-derived scenarios, DPO, LoRA, context distillation, Qwen3.8-27B, and vLLM.
  • Verbosity can confound judging, preference labels, reward, and claims about content-bearing behaviour, motivating length assignment rather than covariate adjustment.The paper uses direct length assignment because length is a post-treatment variable.
  • Prior benchmark studies show evaluation format can alter measured outcomes, including ranking reversals through refusal in free-text generation.The paper extends this concern by testing option-position effects, detector agreement, and answered-set selection.

3 Setup

The setup constructs synthetic Korean style targets and matched controls, trains small-seed LoRA arms, and evaluates deterministic response-style features alongside KoBBQ and securities-disclosure instruments. Measurement and serving choices are fixed or explicitly tracked across arms.

  • 3.1 The paired corpus and the stages: Each corpus record contains a prompt, raw teacher response, style-rewritten target, and style metadata.The style-rewritten target is used for style SFT and DPO, while controls swap which record field becomes the target.
  • 3.1 The paired corpus and the stages: 4,800 synthetic prompts were deduplicated to 4,782, with 226 non-Korean prompts discarded before teacher generation and rule-based rewriting.The resulting style targets follow a light-to-standard-to-heavy ladder and make no human-likeness claim.
  • 3.1 The paired corpus and the stages: The style SFT trained on 519 records after a 2,992-record SFT barely moved the model.The second corpus enforced style axes during generation while holding hyperparameters fixed.
  • 3.1 The paired corpus and the stages: All adapters use LoRA rank 8 with α=16, while style SFT uses one epoch and three seeds; safety DPO uses 2,757 preference pairs.The seed axis jointly changes data order and LoRA A initialization.
  • 3.3 Measurement: Base length is right-censored because 74.6% of responses hit the 700-token cap, so reported gaps against base are lower bounds.An uncensored n=150 rerun yielded a true median of 737 tokens.
  • 3.1 The paired corpus and the stages: The neutral control uses raw teacher responses, while the length-matched control truncates those responses to paired style-target length without rewriting.Prompts were identical across the neutral control and style targets, with three seeds per control arm.
  • 3.1 The paired corpus and the stages: Response style is measured with eleven deterministic features rather than a learned composite score.The suite is independently written from the target-producing rule engine, except for a reused non-Korean-script regex.
  • 3.3 Measurement: Disclosure uses 48 immutable prompts at T=0.2, with a 192-prompt expansion for dose-response analysis.The requirements are drawn from Korean Standard Investment Solicitation Rules.

4 RQ1: the untargeted behaviours move, and the target’s form is why

Matched controls show that response-style alignment moves untargeted answer and disclosure behavior, with answer-rate direction tied to target form rather than fine-tuning on the corpus alone. The length-matched and fourth arms do not isolate which form feature is responsible, while disclosure changes are largely reproduced by assigned response length.

  • 4.2 The sign reverses: The style arms raised ambiguous-answer rates while neutral arms lowered them, and their observed seed ranges did not overlap.The mean gap was 2.343 pp; template-clustered style-versus-neutral intervals excluded zero in all nine cross-pairings.
  • 4.2 The sign reverses: These controls attribute the direction to training-target form rather than fine-tuning on the corpus alone.Neutral, length-matched and style arms were ordered by their mean effects under matched prompts, recipe, data volume, holdout split and serving configuration.
  • 4.3 Which property of the form, we do not know: The length-matched arm fell between style and neutral, ruling out either original target form as a complete explanation.Its mean was −0.487 pp, with separations of 1.040 pp from neutral and 1.302 pp from style.
  • 4.3 Which property of the form, we do not know: The length-matched intervention changed discourse as well as length, so it cannot separate target length from the broader style rewrite.Clarifying questions, information requests, hedges and enumerations all fell toward style-arm rates after truncation.
  • 4.3 Which property of the form, we do not know: The fourth arm was unstable across seeds and therefore did not identify the responsible feature of the form.Its seed range was 4.23 pp, and every pairwise contrast involving it overlapped.
  • 4.3 Which property of the form, we do not know: The design cannot estimate length as a dose because the observed length range is dominated by a gap between arms.The length-change/answer-rate correlation was r = −0.965, but 93.1% of the x-axis span was one between-cluster gap.
  • 4.4 The second axis moved too: Disclosure rates are detector-identified emissions rather than calibrated compliance estimates, and detector agreement varies by checkpoint.Conditional agreement between the two detectors ranged from 0.44 to 0.99.

5 RQ2: absolute exposure changes are dominated by answer propensity

Across training contrasts, changes in absolute stereotyped exposure are dominated by answer propensity, while conditional-composition terms remain small; disclosure changes are likewise largely reproduced by assigned response length.

  • 5.1 The decomposition is an identity; the allocation is the result: The answer-propensity term ranges from −13.42 to +1.47 pp, while the composition term stays between +0.00 and +0.45 pp across seventeen training contrasts.All seventeen composition terms are positive and below half a percentage point.
  • 5.1 The decomposition is an identity; the allocation is the result: The 93–116% answer-propensity share is reported only for thirteen contrasts whose absolute-exposure differences resolve under template clustering.For unresolved contrasts, the denominator collapses rather than the composition term becoming large.
  • 5.1 The decomposition is an identity; the allocation is the result: Neutral training arms also reduce absolute stereotyped exposure, with the movement again dominated by answer propensity.The allocation therefore covers both increases and decreases produced by different training targets.
  • 5.2 The magnitude is a range: Three style seeds raise absolute stereotyped exposure by 0.57–1.24 pp, but the range reflects seed variation rather than a stable estimate across training runs.The identical-item-set base-to-style contrast is +1.143 pp, with +1.133 pp from answer propensity and +0.010 pp from composition.
  • 5.3 On the compliance axis, the emission variable is length: The style-aligned checkpoint reduces detector-identified disclosure by about 22 pp, while assigned length reproduces most of that gap and leaves 3–5 pp at longer targets.The result establishes sufficiency of response length for reproducing most of the gap, not mediation of the original change.
  • 5.4 Disambiguated-item accuracy as a task-local control: Disambiguated-item accuracy ranges from 86.24% to 89.82% across checkpoints versus 88.20% for base, providing a task-local control.The paper reports this descriptively and does not generalize it beyond KoBBQ’s evidence-given condition.

6 RQ3: evaluation changes with the output distribution

The paper shows that output-distribution changes can invalidate both conditional bias comparisons and disclosure-detector comparisons. Selection into answered items and treatment-dependent detector agreement make these evaluation quantities format-sensitive.

  • 6.1 Conditional-share contrasts are selection-confounded: Conditional stereotyped-share contrasts do not identify changes in content preference when treatment changes which items receive answers.The answered subset is post-treatment, so the comparison is selection-confounded rather than merely affected by parsing.
  • 6.1 Conditional-share contrasts are selection-confounded: 827 of 1,397 stereotyped answers disappear from the style-to-safety transition, while 517 survivors comprise 96.6% of the safety arm’s stereotyped answers.The share rises mainly because the intervention removes denominator items, leaving a concentrated subset of previously stereotyped answers.
  • 6.2 Disclosure-detector agreement is not invariant across treatments: Conditional detector agreement ranges from 0.438 on the style-only arm to 0.986 on the distilled arm, a 54.9-point spread without reference labels.The two independently specified detectors therefore yield treatment-dependent measurements of the same disclosure construct.
  • 6.2 Disclosure-detector agreement is not invariant across treatments: Within-arm detector agreement is flat in length for base and style, but rises only for the distilled arm, so pooled length trends can be mixing artefacts.The within-arm slopes are +0.06 for base, −0.05 for style, and +0.79 for distilled checkpoints.
  • 6.2 Disclosure-detector agreement is not invariant across treatments: The 18-row detector fixture is a unit test rather than validation, so the disclosure result relies on the joint detector distribution rather than accuracy against human labels.Neither detector has been validated against real-output semantic judgments.
  • 6.3 The pre-specified abstention verdict is format-sensitive: A cyclic rotation panel separates option position from item difficulty, while ten salted-hash rotations produce abstention from 92.60% to 93.16%.The pre-specified 93% criterion fails under eight rotations and passes under two, although conditional stereotyped-share direction remains stable.

7 RQ4: composition and order

The two repair orderings compose partially, but their observed differences remain descriptive because each ordering has one run. Additional effects include format-breaking responses and non-Korean-script contamination.

  • 7 RQ4: composition and order: M4’s strict and loose disclosure rates exceed FIN1’s by 1.10 and 0.88 pp, but neither difference is resolved under template clustering.On the shared 48-prompt set, M4 rates are 0.956 strict and 0.961 loose versus FIN1’s 0.945 and 0.952.
  • 7 RQ4: composition and order: The two orderings differ by 4.83 pp in absolute stereotyped exposure, exceeding the observed DPO seed spread by 2.5–3.2×.The comparison supports clearly different resulting models, not stage order as an established cause.
  • 7 RQ4: composition and order: Reseeding changes KoBBQ parse failure from 1.8% to 7.2% and 9.3%, whereas same-session remeasurement moves only from 148 to 150 items.The answered set is primarily a checkpoint property rather than a serving-harness property.
  • 7 RQ4: composition and order: On unrelated held-out prompts, non-Korean-script contamination rises from 4.87% in the style arm to 20.72% after safety training.The increase is attributed to simplified Chinese fragments, while Latin contamination does not move materially.
  • 7 RQ4: composition and order: The safety arm makes 1.82% of ambiguous KoBBQ items unparseable, compared with zero for base and style, before falling to 0.21% after compliance training.The exclusions reflect prose that breaks the response format, not Sino-Korean characters in the scoring window.

8 Limitations

The paper’s evidence is limited by narrow model and benchmark coverage, small training-arm sample sizes, unresolved style mechanisms, and unvalidated detector semantics. Several reported effects are therefore descriptive and scope-bound.

  • 8 Limitations: The bias-axis result is limited to one benchmark family, one language, and one base model, with English BBQ underpowered for the same decomposition.The paper does not establish that the allocation result generalizes to other benchmarks, languages, model families, or disambiguated items.
  • 8 Limitations: Every training arm has at most three runs, and the order comparison has one run per ordering, limiting permutation-test resolution and causal attribution.At three-versus-three, the exact two-sided permutation p-value cannot be below 0.100.
  • 8 Limitations: The study does not identify which individual style feature drives the bias-axis direction because length, rewriting, hedging, question structure, and conditioning remain entangled.The fourth arm intended to separate these factors is unstable across seeds.
  • 8 Limitations: The small composition term is descriptive rather than universal: another treatment could move it substantially, and ratios become unstable when absolute exposure changes are near zero.Four of seventeen training contrasts have intervals covering zero, and percentage points are therefore the primary presentation.
  • 8 Limitations: Assigning response length reproduces most of the disclosure effect, but this bounds mediation from one side and does not establish length as the causal path.The dose-response design supports sufficiency for the observed effect, not a causal mechanism.
  • 8 Limitations: Neither disclosure detector has been validated against human semantic judgments, so detector-positive rates are not calibrated compliance estimates.Claims on this axis are restricted to detector-measured emission, reproduced intervention effects, and cross-arm detector non-invariance.
  • 8 Limitations: The training corpus is not released, although the weights, evaluation harnesses, and corpus-construction description are available.The method is intended to be reproducible without releasing the training file itself.

9 Conclusion

Response-style alignment changed untargeted bias and disclosure behaviours primarily through answer frequency and response length. The findings also show that evaluation instruments must be tested for selection confounding and treatment-dependent measurement agreement.

  • 9 Conclusion: Style alignment changed two untargeted behaviours, with deployment-visible effects carried mainly by how often and how much the model emits.The design does not identify whether latent conditional preferences changed.
  • 9 Conclusion: Across seventeen training contrasts, the answer-propensity term moves by up to 13.4 pp while the composition term stays between +0.00 and +0.45 pp.This decomposition describes where a deployment-visible quantity moved, not latent preference.
  • 9 Conclusion: Conditional-share contrasts are selection-confounded when treatment changes the answered set, and surface-form detectors cannot be assumed invariant across treatments.Both issues are measurable before trusting the corresponding evaluation conclusions.

A The fourth control arm

The short-and-hedging control arm was unresolved because its seeds disagreed, and the decision rule was too weak for its dispersion. Its mean nearly matched the length-matched arm, suggesting the added control did not resolve which form feature mattered.

  • A The fourth control arm: The short-and-hedging arm was unresolved: its three seeds classified as additive, style, and neutral under the prespecified rule.Re-evaluating the base on the serving configuration shifted all arms by a common constant without changing the classification.
  • A The fourth control arm: The decision rule was mis-specified because roughly 1.2 pp decision bands were narrower than the arm’s 2.12 pp per-seed standard deviation.A three-seed interval rule would instead have reported the additive band and flagged low power.
  • A The fourth control arm: The arm had the widest answer-count dispersion, spanning 344 of 8,139 items, and the highest held-out loss at 1.24–1.26.The style, length-matched, and neutral arms spanned 72, 70, and 17 items, with held-out losses of 0.88–0.90, 0.76–0.78, and about 0.74, respectively.
  • A The fourth control arm: The fourth arm’s mean was −0.483 pp, only 0.004 pp from the length-matched arm’s −0.487 pp.Because the common base shift leaves this gap unchanged, both arms can be read as intermediate rather than style or neutral.

B Instrument defects behind body claims

The audit found instrument defects affecting benchmark layout, seed-dependent item selection, and quantitative claims, with consequences carried into the relevant analyses. These defects were treated as results or limitations rather than silently ignored.

  • B Instrument defects behind body claims: The audit checked whether each instrument matched the prose-described measurement and was applied uniformly across compared arms.Contrasts affected by a different scored set were classified as overlapping and tested accordingly; Table 8 lists defects tied directly to main-text claims.
  • B Instrument defects behind body claims: The supposed answer-option-layout flag actually shuffled the item pool before per-category caps, making 73.9% of scored items seed-dependent.The stored layouts reproduced exactly under md5(sample_id) % k, while six of twelve categories were capped at 2,000 items.

Reproducibility and provenance

The study releases weights, model cards, and evaluation harnesses while withholding the paired training corpus. Its measurement ledger records the checkpoint, serving setup, item set, and analysis script for reported quantities.

  • Reproducibility and provenance: Weights, model cards, and evaluation harnesses are released, but the paired training corpus is not.Reported quantities resolve to a ledger recording the checkpoint, cluster, serving configuration, item set, and analysis script.
Loading 2609.11291v1…