NeFut Logo NeFut
中 Admin Login

[CS.AI] Learning What to Forget: Distributional Unlearning for LLM Representation Spaces

Published at: 2026-10-01 22:00 Last updated: 2026-10-06 12:11
#AI #Machine Learning #LLM

Machine learning systems increasingly need to erase the influence of entire data domains—such as toxic language, harmful behavior, or specific topics—rather than individual records.

This problem is formalized as distributional unlearning: selecting a subset from the forget domain so that its removal pushes the training distribution away from an unwanted population while staying close to the desired one. Existing analyses often rely on parametric assumptions that do not suit high‑dimensional LLM representations.

We introduce the Mamushi framework for non‑parametric distributional unlearning. It ranks forget examples with a probabilistic classifier whose Bayes‑optimal logit equals the forget‑to‑retain log‑density ratio (up to an additive class‑prior constant):$$\text{logit}(x)=\log\frac{p_{\text{forget}}(x)}{p_{\text{retain}}(x)}+\log\frac{\pi_f}{\pi_r}$$

Thresholding this population log‑density ratio yields the optimal fixed‑budget selection rule for the removal‑preservation objective. We also provide a non‑asymptotic transfer guarantee that links score‑estimation and threshold‑calibration errors to the degradation from the population‑optimal rule.

Empirical evaluation on real‑world datasets for toxic‑language removal and topical‑domain removal, using various representations, shows that Mamushi achieves a more favorable removal‑preservation trade‑off than existing baselines and reduces the number of forget examples needed to meet a fixed forgetting target.

Review

Original Source: https://arxiv.org/abs/2609.38929

[h] Back to Home