Offline multi‑agent reinforcement learning (MARL) learns cooperative policies from fixed datasets and the learned policy is frozen at deployment. The frozen policy proposes a single joint action and executes it directly, which often leads to a sub‑optimal choice even when better nearby alternatives exist in the behavior data.
To tackle this issue, we introduce Gradient Guided Multi‑Agent Flow (G2MAF), a test‑time refinement framework for joint policies. G2MAF applies one globally normalized, projected critic gradient to guide and coordinate corrections of all agents while keeping the corrected action feasible and close to the original frozen proposal.
Across 24 MPE and SMAC environments, the canonical variant improves 20 frozen settings, achieving mean relative gains of 9.2% on MPE and 8.9% on SMAC, with model inference latency increased by only about 6%.
The results demonstrate that test‑time gradient guidance can substantially boost the performance of offline MARL deployments with minimal computational overhead.
Review