NeFut Logo NeFut
中 Admin Login

[CS.AI] A Mechanistic Study of AI-Text Detection Neurons in Frozen BERT: Sparse Probing and Activation Patching

Published at: 2026-09-29 22:00 Last updated: 2026-09-30 01:41
#AI #Machine Learning #LLM

This work investigates which neurons in a frozen BERT‑base‑uncased encoder are responsible for AI‑generated text detection. Using the RAID benchmark across six generators (both pure‑base and instruction‑tuned), we apply the L1‑to‑L2 sparse‑probing protocol of Gurnee et al. (2023) to all 9,216 CLS hidden‑state dimensions (12 layers × 768), referring to them as neurons.

We find a stable set of fewer than 1% of neurons per generator, consistent across folds and random seeds. A probe limited to this set retains most of the full‑feature detection accuracy. Bidirectional activation patching confirms the causal relevance: predictions flip an order of magnitude more often than when using size‑matched random sets.

Mean‑ablating the same neurons leaves overall accuracy largely unchanged, indicating that the signal is redundantly distributed. Cross‑generator analysis reveals a bipartite pattern: instruction‑tuned generators concentrate 30‑36% of stable neurons in BERT’s final layer, whereas base generators stay below 14%, consistent with a layer‑12 footprint of post‑training alignment.

Leave‑one‑family‑out evaluation shows that the selected neurons preserve 86‑94% of the full‑feature ceiling on unseen generator families, suggesting that a detector can operate within a small fixed subspace without re‑identifying neurons for each generator.

Review

Original Source: https://arxiv.org/abs/2609.30287

[h] Back to Home