Subliminal learning enables language models to acquire behavioral traits from training data that lack an obvious semantic link to those traits, thereby weakening safety measures that rely on content‑based filtering. Data attribution aims to pinpoint the training examples responsible for a given model behavior, independent of semantic content, and could therefore work where semantic inspection fails. We evaluate three gradient‑based attribution methods—GradCos, a contrastive GradCos variant, and EK‑FAC—across three models and compare them to divergence tokens, a strong baseline that requires access to counterfactual teacher models.
At the token level, EK‑FAC mitigates a substantial portion of the subliminal effect, while the other methods provide little benefit and all fall short of divergence tokens. Filtering whole samples is less effective for every method, although EK‑FAC often yields a stronger signal than divergence tokens in this setting. Success varies inconsistently across methods and model‑preference combinations, and we do not find a unified explanation for these differences.
Our results suggest that gradient‑based attribution can identify data responsible for subliminal learning in certain scenarios, but the reliability of the approximations differs, and no method works universally.
Review