In the Corr2Cause task, the goal is to decide whether a causal claim holds in every DAG that is compatible with the observed correlations and conditional independencies. We cast this as latent‑object reasoning: the label is defined by a CPDAG query, yet free‑form chain‑of‑thought often collapses the Markov‑equivalence‑class problem into local pattern matching. To address this, we propose Structured Thinking, a two‑turn pipeline: the first turn externalizes a typed, schema‑constrained CPDAG summary; the second turn answers the query using that graph state.
On the full Corr2Cause test set, Structured Thinking raises $F_1$(Yes) for Qwen3.5-27B from $73.0$ to $86.4$, a $13.4$‑point gain over a strong PC‑instruction baseline (McNemar $p=2.4\times10^{-6}$, bootstrap $95\%$ CI [$+8.4$, $+18.6$]). Across three full‑ID seeds the mean improvement is $+8.1 \pm 5.3$ points. A two‑turn prose control that only provides a detailed PC scaffold reaches $67.6$ $F_1$, indicating that a schema‑free intermediate is insufficient. The same pattern repeats on Qwen3.6-27B, Paraphrase‑OOD, and GPT‑5.4‑mini. Scrambling the emitted CPDAG costs $12.0$ points, and a full‑split audit shows close agreement with the reference CPDAG (ID skeleton $F_1$ $0.960$, exact match $75.9\%$). These findings support a bounded design principle: externalize the latent object that defines the label, constrain its form, and test whether downstream answers truly use it.
Review: Structured Thinking’s explicit causal graph representation yields substantial performance gains, confirming the value of externalizing latent objects for LLM reasoning.