We investigate how a language model’s reliance on query‑routing signals and target knowledge evolves while answering a question. By applying layer‑wise interventions to the hidden state at the end of the question, we compare country‑continent prompts across Qwen, Llama and Gemma, using noun, adjective and code answers while keeping several fitted metrics distinct.
A pair‑conditioned request direction denotes the country naturally queried in single‑country questions; a global request direction distinguishes first‑ versus second‑country requests in paired prompts; separate selection candidates serve as controls for content already present in the hidden state. Re‑analysis of frozen Qwen natural‑question states shows that the pair‑conditioned direction strengthens before interventions begin to alter later fitted knowledge, and this causal window opens while answer‑supporting content is still forming.
The three‑model trajectories are not uniform: Gemma exhibits a partially overlapping mid‑layer routing‑content profile, whereas Llama shows no sustained routing‑effect window under the same gates. In the paired protocol, dependence on the global request direction declines from early to later layers, while dependence on fitted content persists. A matched Qwen comparison reveals that the pair‑conditioned direction retains a late effect, indicating that the operational handoff concerns the global fitted direction rather than all request information.
These results separate early readability, natural strength, causal steering, and later content dependence.
Review