Structured tool calls often break after only a few fields violate a schema or an execution contract. Regenerating the whole object expands the action space and makes repeated repairs hard to audit. ContractRL addresses this by introducing a contract‑constrained sequential repair protocol that treats verifier‑guided JSON repair as a bounded decision process. At each step the policy receives the candidate JSON, typed verifier feedback, a JSON Pointer, an immutable repair history, and the remaining budget. A contract‑derived action mask filters malformed or prohibited RFC‑6902 operations before a deterministic validator applies the transition. While keeping canonical targets and semantic labels out of the online state until trace freeze, we define a contract‑constrained group‑relative objective for patch, retry, and abstention decisions. Experiments on 192 cases per seed across five seeds show that, under identical verifier information, ContractRL achieves a semantic success rate of 0.9362 with only 34.4 generated tokens, compared to 0.9076 and 44.9 tokens for Patch‑SFT and 0.9148 and 137.2 tokens for full regeneration. Policy optimization improves supervised ContractRL from 0.9186 to 0.9375. A three‑seed paired evaluation against Patch‑SFT yields a semantic gain of $+0.0396$ (95% CI $[+0.0137,+0.0662]$, $p=0.0039$). Analyses of feedback, action‑mask, budget, and schema‑shift link these gains to localized correction, while adversarial and multi‑turn tests expose remaining failure modes.
Review