DNAlign is a lightweight alignment framework that blends control‑theoretic optimization with null‑space projection for large language models. By treating an LLM as a dynamic system, the method injects controllable perturbations during generation to steer outputs toward safe behavior. The projection module extracts a harmful‑related subspace from neutral hidden states and restricts perturbations to this subspace, preserving the model's general knowledge and response quality.
A value function trained on human‑preference data adaptively optimizes the control signals to match human safety preferences. Extensive evaluations on various LLM backbones show that DNAlign consistently reduces harmful generations while maintaining fluency, coherence, and factual accuracy. Compared with prior alignment baselines, it achieves superior overall performance without sacrificing generation diversity, indicating a practical solution for safe LLM deployment.
Review