Machine unlearning aims to erase targeted information while preserving a model's other capabilities. In real‑world scenarios such as GDPR privacy requests, the target can be extremely fine‑grained, e.g., data associated with a single individual. Behavioral forgetting alone is often insufficient, prompting interventions directly on internal representations. Conventional mechanistic‑interpretability extractors are poorly selective for such narrow targets because reconstruction‑based extraction exhibits an energy bias: it favors dominant background structure and overlooks low‑energy, target‑specific components. To address this, we introduce SCALPEL, a contrastive sparse autoencoder that learns more selective forget features. Theoretically, contrastive training amplifies target‑related signals, and our selection score bounds the expected perturbation of background knowledge. Experiments on the TOFU benchmark across Qwen, Llama, and Gemma show that SCALPEL substantially outperforms NMF and standard SAE interventions, while remaining competitive with Gradient Difference and RMU, thereby bridging mechanistic interpretability and fine‑grained unlearning.
Review