NeFut Logo NeFut
中 Admin Login

[CS.AI] Are Stated Reasoning Steps Causally Load‑Bearing?

Published at: 2026-09-24 22:00 Last updated: 2026-09-28 00:49
#AI #Machine Learning #LLM

Chain‑of‑thought (CoT) monitoring assumes that the reasoning a model writes mirrors the computation that directly yields its answer. Prior faithfulness metrics have been largely behavioral: edit the reasoning text and observe the resulting answer. Our approach instead measures faithfulness causally at the activation level, focusing on self‑generated reasoning. Unlike earlier causal audits that only detect degradation, our interventions carry a known predicted target—each patch should switch the answer to a specific counterfactual entity that can be constructed by design. We employ synthetic multi‑hop lookup tasks (2‑6 hops) and patch the residual stream at the token span where the model states each intermediate step, inserting the corresponding activations from a counterfactual run. For Qwen3‑4B, 76.9% ± 2.8% of stated steps are causally load‑bearing (CLB) at the most responsive mid‑network layer (random‑position null: 11.3%; patching the underlying prompt fact: 83%, indicating that stated steps carry roughly 96% of the achievable effect). The standard behavioral test on the same items yields 88.2%, overstating causal faithfulness by 11.4 percentage points (item‑matched; 111:14 discordant pairs, p < 0.05).

Review

Original Source: https://arxiv.org/abs/2609.27038

[h] Back to Home