NeFut Logo NeFut
中 Admin Login

[CS.AI] Backdoor Purification for LoRA-Tuned LLMs via Null-Space Projection

Published at: 2026-10-02 22:00 Last updated: 2026-10-06 12:11
#AI #Machine Learning #LLM

The rapid deployment of large language models (LLMs) together with parameter‑efficient fine‑tuning (PEFT) has amplified the threat of backdoor attacks. Existing purification techniques usually rely on strong assumptions such as known triggers, clean reference data, or aggressive retraining, and they often lack thorough evaluation, limiting real‑world applicability.\ \ This paper introduces a purification method for LoRA‑tuned LLMs that requires none of these assumptions and does not need post‑hoc retraining of the suspect parameters. The goal is to dramatically cut the attack success rate (ASR) while preserving (i) the base model’s general capabilities and (ii) the downstream skills learned via the adapter.\ \ The approach proceeds as follows: high‑fidelity backdoor directions are extracted through careful data curation and feature approximation; for each layer or attention head, orthogonal null spaces are constructed in both input and output channels; LoRA updates are then projected onto these null spaces, effectively nullifying the backdoor component.\ \ A series of ablation studies scale the method from a single layer in a text‑classification setting to a full‑parameter LLM in a generative task. Empirically, the null‑space projection reduces ASR from nearly 100% to below 10% while keeping the base model’s benign performance and the adapter’s downstream abilities essentially intact.\ \ The work demonstrates a practical, mathematically grounded pathway to cleanse LoRA parameters without sacrificing model functionality.\ \ Review

Original Source: https://arxiv.org/abs/2610.00685

[h] Back to Home