NeFut Logo NeFut
中 Admin Login

[CS.AI] Beyond Refusal Patterns: Safe-Role Internalization for Robust and Generalizable LLM Safety Alignment

Published at: 2026-10-07 22:00 Last updated: 2026-10-08 01:25
#AI #Machine Learning #LLM

Large Language Models (LLMs) have achieved impressive generation capabilities but remain vulnerable to jailbreak attacks that provoke harmful or unsafe outputs. Existing safety alignment methods such as Supervised Fine‑Tuning (SFT) and Reinforcement Learning from Human Feedback (RLHF) often require extensive attack‑specific supervision and computational cost, yet they still suffer from shallow alignment and over‑refusal.\ \ To overcome these limitations, we introduce SSRFT (Supervised Safe‑Role Fine‑Tuning), the first framework that reformulates safety alignment as the internalization of a predefined safe role. SSRFT builds a Safe‑Role Question‑Answer (SRQA) dataset from psychometric questions, a limited set of jailbreak prompts, and a safe‑role description. Role‑consistent answers are first synthesized, then validated, and finally expanded into diverse scenarios, enabling the model to absorb safety‑oriented values and principles rather than merely learning refusal patterns.\ \ Experiments are conducted on several Base and Instruct models, comparing SSRFT with standard SFT. Results show that SSRFT achieves markedly higher robustness against prefix‑injection attacks, better generalization to unseen jailbreak domains, and a substantial reduction in over‑refusal on benign queries, all while preserving the model’s general capabilities.\ \ These findings establish safe‑role internalization as an effective alternative to refusal‑centric safety alignment, offering a promising direction for building more reliable LLMs.\ \ Review: SSRFT’s role‑internalization approach delivers robust and generalizable safety alignment, meriting further validation and adoption in real‑world deployments.

Original Source: https://arxiv.org/abs/2610.07023

[h] Back to Home