Platform operators are increasingly relying on system prompts and fine‑tuning to steer model behavior, yet it remains unclear how reliably these interventions can override behavior inherited from prior training. We introduce Override Success Rate (OSR) and alignment inertia to quantify when interventions succeed or fail. OSR is defined as $$\text{OSR}=\frac{\text{successful overrides}}{\text{total interventions}}$$, while alignment inertia captures the tendency of a model to retain its original behavior after an intervention.
We evaluate zero‑shot prompting and LoRA fine‑tuning on Llama and Mistral models across two domains: medical misinformation and hate speech. Results show that alignment inertia persists in both models but varies markedly with model architecture, domain, and policy direction. Notably, in Mistral’s restrictive hate‑speech setting, LoRA fine‑tuning increased inertia by 46.5 percentage points, indicating that fine‑tuning can sometimes reinforce rather than override prior behavior.
We also employ the TRAK method to test whether inertia correlates with weaker adaptation signals. TRAK achieves an AUC of at least 0.85 in 7 out of 8 conditions, outperforming baselines such as model confidence, TF‑IDF similarity, and embedding similarity. These findings provide operators with an audit tool to pinpoint where prior training constrains downstream model governance.
Review