NeFut Logo NeFut
中 Admin Login

[CS.AI] Hallucination Neurons and Where to Find Them: An Investigation into the Existence of Hallucination Neurons

Published at: 2026-09-26 22:00 Last updated: 2026-09-28 00:49
#Machine Learning #LLM #Artificial Intelligence

Interpretability methods for large language models (LLMs) increasingly rely on sparse probing techniques that isolate small sets of neurons purported to detect and causally influence behaviors such as factual recall, safety alignment, and hallucination. These claims carry significant implications for model auditing and behavioral steering, yet they are seldom evaluated against known failure modes of $L_1$‑regularized probing in correlated, high‑dimensional feature spaces.\

We introduce a five‑step diagnostic protocol covering feature correlation, bootstrap stability, disagreement between sparse and dense rankings, intervention baselines, and cross‑dataset evaluation as a minimum standard for sparse‑neuron localization claims. Applying this protocol, we re‑examine prior work on H‑neurons using open‑source LLMs across the TriviaQA, BioASQ, and NQ‑Open datasets. Our findings show that detection replicates across models and datasets and surpasses the originally reported AUROC gaps for TriviaQA and BioASQ. Gemma 3 4B consistently outperforms MedGemma 4B on matched datasets, with AUROC gaps of +0.311 vs. +0.235 on TriviaQA, +0.474 vs. +0.455 on BioASQ, and +0.128 vs. +0.112 on NQ‑Open.\

Causal validation at $n = 500$ with five random seeds demonstrates statistically significant effects beyond random same‑layer baselines. At the same time, diagnostic results indicate that the selected neurons are not uniquely localized: across three Gemma 3 4B settings, 19 of 22 H‑neurons exhibit Pearson $|r|$ > 0.7 with other features, bootstrap selections show only moderate stability, and sparse and dense rankings overlap weakly. Our results reveal that sparse predictive structure can coexist with non‑unique neuron selection.\

Consequently, routine diagnostic validation is essential to distinguish detection claims from localization claims in mechanistic interpretability research, preventing misleading conclusions.\

Review

Original Source: https://arxiv.org/abs/2609.29781

[h] Back to Home