We post‑train Qwen3.8‑27B to adopt a Korean response style—verbosity, use of lists and Markdown, full discourse structure, and a specific register—and evaluate two behaviours that are not directly targeted by the training objective.
The first behaviour is abstention on ambiguous social questions in KoBBQ (benchmark answer UNKNOWN); the second is unprompted disclosure in securities guidance. Both shifts are expressed mainly through the model’s emission policy: how often it answers and how many tokens it emits.
Matched target‑form controls reveal that answer propensity depends on the training target itself, not on the prompt set or fine‑tuning recipe. With prompts, recipe, data volume, and serving fixed, only the target text varied: three “style seeds” increased answer rate by about $+0.82\text{pp}$, while three “neutral seeds” decreased it by about $-1.53\text{pp}$; the ranges do not overlap and the means differ by $2.34\text{pp}$. A length‑matched arm falls between them, and a fourth arm that stays short while preserving hedging is unstable across seeds, leaving the exact form feature responsible unresolved.
Decomposing the overall change into an answer‑propensity term and a conditional‑composition term is merely an algebraic identity; the empirical content lies in where the movement occurs. Across checkpoints the change is dominated by answer propensity, with the composition term remaining small, and because the latter is evaluated on treatment‑dependent answered subsets it does not constitute evidence of a latent preference.
Two measurement findings follow: (1) when answer status is treatment‑dependent, a between‑arm contrast in conditional stereotyped share does not identify a change in content preference; (2) agreement between two rule detectors for the same construct varies from $0.44$ to $0.99$ across checkpoints, observable without any reference labels.
Review