NeFut Logo NeFut
Admin Login

[CS.AI] Innocuous Data, Latent Ideology: Ideological Generalization in Finetuned LLMs

Published at: 2026-07-18 22:00 Last updated: 2026-07-22 01:03
#algorithm #AI #Machine Learning

Finetuning language models on small, curated datasets is standard practice for adapting them to specific policies or domains. We demonstrate that finetuning on narrow, factually-defensible, moderation-passing data can lead to broad ideological shifts across unrelated domains while preserving general capabilities.

Training GPT-4.1 on right- or left-leaning economics Q&A results in matched ideological shifts on topics such as criminal justice, the environment, and cultural taste. This effect also appears with plausibly-deployed datasets like workplace HR policy and practical finance queries, as well as on a science-pseudoscience axis where food-safety finetuning increases sycophantic agreement with users expressing false health beliefs.

We term this phenomenon ideological generalization and propose a methodology to measure two properties: breadth, indicating how far the shift extends across topics absent from training, and amplification, measuring how much finetuning intensifies the shift relative to few-shot prompting on the same examples. We show that few-shot prompting indicates the direction of generalization, but finetuning pushes the model to further extremes, including endorsements of race-IQ connections and political violence.

The effect replicates on Gemma-3, holds under judge-free evaluations and external benchmarks, and survives mixing with generic data, maintaining GSM8K accuracy within ±1pp of the baseline.

Blogger's Review: This study reveals the potential ideological biases that can arise from finetuning language models on specific datasets, highlighting the need for caution in the use and development of LLMs, especially concerning socially sensitive topics. A deeper understanding of ideological generalization can aid in designing models that mitigate potential bias risks.

Original Source: https://arxiv.org/abs/2607.14888

[h] Back to Home