NeFut Logo NeFut
中 Admin Login

[CS.AI] Does the Model Use the Feature? Separating Steering from Mechanism in LLMs

Published at: 2026-10-07 22:00 Last updated: 2026-10-08 01:25
#AI #Machine Learning #LLM

Internal features in large language models (LLMs) are often taken as mechanisms: when a feature tracks a concept and manipulating it changes a related behavior, we tend to infer that the model uses that feature. However, steering can push a feature far beyond its natural range, and its effects may no longer reflect the model’s own computation. To test this inference, we propose an empirical contract that evaluates features only at values observed on natural inputs.

The contract defines three interventions:

Installation measures how far the feature suffices for the behavior; removal and downstream rescue measure how much the model actually uses the feature. We apply this framework to three kinds of representations: an unknown‑entity latent, dense known‑unknown directions, and a released subject‑verb agreement feature set. The experiments reveal sharp separations between the two strengths (installation vs. usage).

Thus, tracking a concept and steering a behavior alone do not demonstrate that the model uses the feature; each conclusion holds only for the specific intervention tested.

Review

Original Source: https://arxiv.org/abs/2610.07270

[h] Back to Home