Internal features in large language models (LLMs) are often taken as mechanisms: when a feature tracks a concept and manipulating it changes a related behavior, we tend to infer that the model uses that feature. However, steering can push a feature far beyond its natural range, and its effects may no longer reflect the model’s own computation. To test this inference, we propose an empirical contract that evaluates features only at values observed on natural inputs.
The contract defines three interventions:
- Installation: copy a feature’s value from an input that exhibits a behavior into a matched input that does not.
- Removal: replace the feature value in a behavior‑showing input with the value from a non‑behaving counterpart.
- Downstream rescue: after an upstream edit, restore the feature value and observe whether the behavior returns.
Installation measures how far the feature suffices for the behavior; removal and downstream rescue measure how much the model actually uses the feature. We apply this framework to three kinds of representations: an unknown‑entity latent, dense known‑unknown directions, and a released subject‑verb agreement feature set. The experiments reveal sharp separations between the two strengths (installation vs. usage).
- The unknown‑entity latent strongly steers knowledge abstention, yet installing observed latent values into matched prompts transfers only a tiny fraction of the natural known‑unknown abstention contrast.
- Dense known‑unknown directions exhibit opposite asymmetries between installation and removal in Gemma and Llama, indicating that the same feature is leveraged differently across architectures.
- For the released subject‑verb agreement feature set, the extent to which behavior is reproduced and restored depends on how the feature values are written into the model; different writing strategies lead to markedly different rescue outcomes.
Thus, tracking a concept and steering a behavior alone do not demonstrate that the model uses the feature; each conclusion holds only for the specific intervention tested.
Review