A team of code‑generating LLM agents may pass each individual test yet produce a broken codebase when their outputs are merged. Conflicts arise when two agents rewrite the same function, or when one writes against an interface that a teammate has just changed, leading to integration failure. Existing coordination tools typically detect conflicts after they happen, which is too late at agent speed.
We reformulate the issue as a scheduling problem: extract the declared scope of each work item, partition the work into disjoint scopes, and pre‑order merges according to the producer‑consumer graph. This planner is integrated into Nerveplane and evaluated with NP‑Bench, a three‑arm benchmark (no coordination, reactive detection, proactive planning) that validates integration on a real Git merge, both in deterministic simulation and with live agents.
Results show the planner lifts clean‑integration from 1/9 to 9/9 and reduces merge conflicts from 13 to 0, with the gap widening as the number of agents grows. In a live breaking‑contract scenario, the planner rescues outcomes missed by both baselines for every seed: clean‑integration rises from 0 to 1.0 on a frontier model and to 0.6 on a smaller model, while agents respect assigned scopes (0/5 leakage). Cross‑session memory drops the repeated‑mistake rate from 1.00 to 0.00 for both strong and weak models.
We also report a negative finding: routing factual context to agents does not improve long‑context accuracy at window‑fitting scales; its benefit lies in cost and capacity rather than attention. This advantage persists across two capability tiers and two vendors because it stems from how work is allocated, not from model reasoning.
Review