NeFut Logo NeFut
中 Admin Login

[CS.AI] Aligned Data Can Induce Misalignment via Context Confusion

Published at: 2026-10-01 22:00 Last updated: 2026-10-06 12:11
#AI #Machine Learning #LLM

Large language models (LLMs) are often updated for various use cases, and training pipelines typically filter out samples that could cause misalignment after the update. Alignment, however, depends on context: a recommendation that fits one scenario may be unsuitable in another. For instance, advising a researcher to preserve data for reproducibility is aligned, while the same advice to a mobile‑app developer about users' sensitive data may breach privacy. From this observation we identify a post‑training phenomenon where aligned fine‑tuning induces misaligned behavior in different contexts, which we call context confusion. We demonstrate this effect in three domains—gender equality, privacy, and physical safety. Our results show that context confusion leads to narrow misalignment rather than emergent misalignment; adding generic alignment data does not substantially help, whereas targeted alignment data for the problematic domain or providing in‑context examples at inference time can markedly reduce the issue. Mechanistically, queries from distinct domains undergo similar representational shifts during fine‑tuning, causing a query from another domain to trigger the same learned behavioral feature and produce misaligned output. These findings suggest that inspecting training data alone is insufficient to predict a model’s alignment state, highlighting the need for comprehensive post‑training alignment evaluations.

Review

Original Source: https://arxiv.org/abs/2609.38379

[h] Back to Home