A transformer can make an attribute linearly decodable in its residual stream at a depth where that attribute does not yet influence the output. This gap between readability and causal use has been shown for attributes directly present in the input. We ask whether it also holds for an attribute that must be inferred gradually over a conversation—namely, how expert the dialogue partner is.
We employ ExpertCollab, a corpus of multi‑turn research‑planning dialogues between model‑played personas at four expertise levels. The results show that partner expertise is most decodable in the early layers and drops to near chance before the network’s midpoint.
Counterfactual patching reveals that injecting the expertise difference at the layer of peak decodability barely changes a fixed late‑layer readout, whereas the same injection past the midpoint propagates almost completely, a separation of more than an order of magnitude.
A content‑matched random control and a probe‑free diagnostic place the transition at the same early layer, while a statically specified control attribute remains decodable throughout.
Thus, an inferred relational attribute is represented well before it becomes causally active, bounding where any attempt to read out or steer partner‑conditioned behavior must intervene. We demonstrate an initial proof‑of‑concept on a synthetic corpus using a single model.
Review