Large language models operating in multilingual settings must decide the target language at the very start of generation, yet the causal circuitry that governs first‑token language identity remains poorly mapped. We conduct an end‑to‑end structural circuit analysis on six architectures—GPT‑2, BLOOM‑560M, Pythia‑1B/2.8B, and Qwen2.5‑1.5B Base/Instruct. First, Edge Attribution Patching (EAP) is applied with FP16 activation clamping, followed by exact activation‑patch verification limited to a 2,000‑candidate‑edge search, to extract directed acyclic graphs that drive first‑token language broadcasting.
Across the standalone models we find deep or mid‑to‑deep broadcasting hubs; the strongest evidence appears for Pythia‑2.8B and BLOOM‑560M, while GPT‑2 and Pythia‑1B provide few out‑of‑graph heads for comparison, and both Qwen2.5‑1.5B variants invert the necessity check. Scaling from Pythia‑1B to 2.8B expands node participation yet keeps the verified edge budget similar, yielding a sparser topology. The Qwen2.5‑1.5B base and instruct circuits share 84.7% Jaccard similarity, including the Layer 27 hub, indicating that first‑token routing is largely established during pre‑training and preserved after instruction tuning.
Moreover, EAP scores correlate only weakly with exact patching deltas for most models, showing that linear gradient approximations can diverge from causal interventions in FP16 and motivating exact‑patch verification for reliable circuit discovery.
Review