Cooperative multi-agent tasks require different agents to execute joint actions simultaneously, and each agent's action influences the observations and responses of the others. Consequently, a world model is needed to predict the team return generated by the joint actions. A naive approach that applies a single-agent world model to each agent step‑by‑step fails to capture dependencies among simultaneous actions. To address this, we introduce MA-WAM (Multi-Agent World-Action Model), a test‑time planning framework that lets a frozen multi-agent flow policy evaluate the futures of candidate joint actions. MA-WAM is the first test‑time world‑model planner for multi-agent flow policies. It predicts the consequences of each joint action based on cross‑agent dependencies and enables efficient candidate scoring. Experiments on 30 offline MARL settings across MAMuJoCo, SMAC, and MPE show that MA-WAM achieves an average improvement of 22.0% over direct execution and 25.6% over uniform action selection. Under the standard evaluation protocol on an A100 GPU, MA-WAM adds 12.1 ms, accounting for only 2.5% of the total generation‑and‑scoring time.
Review