Long‑horizon agentic tasks require strong reasoning and efficient execution across successive interactions with dynamic environments. A typical solution separates high‑level planning from low‑level execution, assigning distinct planner and actor modules. To study coordination failures, we prompt both modules to emit structured state assertions and programmatically compare them to spot explicit contradictions. Our analysis shows systematic disagreement on the same task‑relevant facts, a phenomenon we call planner‑actor state mismatch. Experiments reveal that supplying task‑relevant state information to both modules reduces the mismatch and improves coordination and success rates. Building on this systematic analysis, we introduce Consistent Plan‑Act (ConPAct), which feeds detected contradictions back to the planner and actor to form a consistent state interpretation and fine‑tunes them on curated consistent interactions for better coordination. ConPAct boosts performance across various environments and model configurations; for instance, in MiniGrid the success rate rises from 38.6% to 54.4% when using GPT‑5.6‑sol as planner and terra as actor. These findings demonstrate that state consistency can guide both inference‑time correction and coordination training.
Review