GPS-Bench is an evidence‑grounded benchmark for governance policy simulation that links policies to relevant actors, their actions, and downstream impacts using legislative records, lobbying disclosures, regulatory documents, corporate filings, economic data and other public sources. Actors are reconstructed from dated records rather than prompted as archetypes, so each persona is an evidence object with provenance. A human‑annotated pool forms the Gold evaluation set, while a separate LLM generates Silver supervision from retrieved evidence, never used as test labels.
Because every inference mode reads the same grounded state and emits the same schema, GPS‑Bench turns the question “does multi‑agent simulation help?” into a controlled comparison. We contrast joint reasoning, independent yet communicating agents, graph‑based methods, and weight‑level fine‑tuning over a single policy state. Fine‑tuning on the grounded record yields the strongest actor‑level impact prediction; decomposition does not surpass it, but it adds mechanistic insight.
Agents hold private, non‑identical evidence, each seeing only its own exposure clause, and address named partners with concrete joint proposals, what they offer, what they need in return, and why acting together beats acting alone. The resulting coalitions can be checked against the commitments recorded in the evidence, providing a verification layer.
Thus, GPS‑Bench offers a common empirical setting for studying when evidence, actor modelling, and multi‑agent interaction improve the prediction and interpretation of policy outcomes.
Review