The emergence of cooperation is a central challenge in multi‑agent systems, where decentralized agents must coordinate while adapting to the evolving behavior of others. Evolutionary game theory can identify strategically stable outcomes, but stability under a population‑adjustment dynamic does not necessarily imply that finite‑sample learning agents can reach the same outcome through local reward feedback.
We illustrate the distinction with a transparent three‑agent governance game involving a government, a platform firm, and users. First, we derive the replicator dynamics for the fixed stage‑game incentives, expressed as
$$ \dot{x_i}=x_i\big[(A\mathbf{x})_i-\mathbf{x}^T A \mathbf{x}\big] $$
where $x_i$ denotes the proportion of strategy $i$ and $A$ is the payoff matrix. By evaluating the cooperative evolutionary basin on a symmetric initial‑condition grid, we obtain a volume $V_E=1.00$, indicating that cooperation is evolutionarily stable across the sampled space.
We then compare this with learning‑basin estimates for three decentralized value‑based learners under the same payoff environment: independent $\varepsilon$‑greedy Q‑learning ($\varepsilon$‑IQL), scaled Boltzmann exploration, and SA‑EA BQL. The learning basin is measured as the fraction of initial points that converge to the cooperative joint action on the identical grid. Empirically, $\varepsilon$‑IQL achieves a basin of $0.88$, while both scaled Boltzmann and SA‑EA BQL obtain $0.00$.
Diagnostic traces reveal that broader action diversity and non‑zero value separation can coexist with a failure to sustain cooperative joint action in this fixed configuration. These findings demonstrate that evolutionary stability and learning accessibility are distinct properties of a coupled game‑learning system.
The shared‑bike scenario serves as a motivating application; the broader contribution is a framework for comparing population‑level stability with the finite‑sample accessibility of cooperation under specified multi‑agent learning dynamics.
Review