This study investigates how large language models (LLMs) internalize human moral biases during finetuning, focusing particularly on the well-known Knobe effect, which manifests in intentionality judgments. The mechanisms behind this bias remain unclear.
We conducted a Layer-Patching analysis across three open-weight LLMs and found that the bias is not only learned during finetuning but is also localized to specific layers of the model.
Surprisingly, we discovered that patching activations from the corresponding pretrained model into just a few critical layers is sufficient to eliminate the effect. Our findings provide new evidence that social biases in LLMs can be interpreted, localized, and mitigated through targeted interventions without the need for model retraining.
Blogger's Review: This research insightfully reveals the complexities of moral bias in finetuned large language models, emphasizing a new strategy for addressing biases through localized interventions rather than comprehensive retraining. The use of layer patching offers a novel perspective for understanding and improving moral judgments in LLMs, holding significant theoretical and practical implications.