Large language models (LLMs) have been applied to causal discovery, yet candidate‑graph generation rarely treats the premature omission of potentially relevant causal relations as an explicit design objective. We introduce MaSCoD, a multi‑agent framework that first organizes candidate third variables and local structural patterns before performing direct‑edge judgment.
We evaluate MaSCoD on the Auto‑MPG, DWD, and Sachs datasets, using GPT‑5.4 as the primary backbone and GPT‑4o for replication. Results show that MaSCoD exhibits a dataset‑ and backbone‑dependent retention‑selectivity profile rather than uniform superiority.
Across all six dataset‑backbone settings, Full (which supplies structural hypotheses before direct‑edge judgment) achieves higher mean Recall and F1 than No Phase 1 (which constructs them within the judgment procedure), while also increasing false‑positive rates. Additional reference‑edge retention is observed on DWD with GPT‑5.4 and on Sachs with GPT‑4o, rather than uniformly across settings.
Partial ablations reveal that providing both information components does not always outperform supplying only one. Stage‑wise analysis for GPT‑5.4 shows that the Full‑No Phase 1 retention gap already appears after direct‑edge judgment, and reconciliation introduces further reference‑edge loss for Full on Sachs.
These findings support structural pre‑organization as an explicit design and evaluation target for omission control and motivate joint evaluation of context construction and its utilization in judgment.
Review