This paper introduces the Flag Game, a toy model designed to study the mechanisms of collective belief formation. A hidden country flag represents the ground truth; each bounded agent can only directly observe a private crop of the flag but may exchange beliefs with peers and weight social evidence from them.\ \ Despite its simplicity, the model reproduces a rich set of swarm phenomena:\
- Performance scales non‑monotonically with population size;\
- Social‑awareness prompting and team diversity boost accuracy;\
- Organizational structure strongly influences overall outcomes.\ \ A central observation is that small populations experience belief collapse, while larger populations undergo a transition to belief polarization. Polarization explains the performance drop at large scales but also generates diverse collective beliefs.\ \ To dissect these mechanisms, the authors employ two complementary approaches. The first is social circuit attribution, which predicts which agent and which view matter most to the collective dynamics. Causal interventions (patching) on the identified agents confirm the predictions, yet the effectiveness of such interventions diminishes as the population grows. The second approach builds a statistical‑mechanical theory for large populations, deriving a phase diagram that matches empirical observations and clarifies the macroscopic origin of belief polarization.\ \ Together, these results constitute a first step toward mechanistic swarm interpretability—a systematic science of how individual properties and communication rules give rise to emergent collective behavior.\ \ Review: The Flag Game captures essential dynamics of belief propagation in multi‑agent systems with a minimal setup, offering both theoretical insight and experimental validation that can guide the design of interpretable and safe collaborative AI.