NeFut Logo NeFut
中 Admin Login

[CS.AI] Evolutionary Stability Does Not Guarantee Learning Accessibility: A Multi-Agent Reinforcement Learning Perspective on Cooperation Emergence

Published at: 2026-09-24 22:00 Last updated: 2026-09-28 00:49
#algorithm #AI #Machine Learning

The emergence of cooperation is a central challenge in multi‑agent systems, where decentralized agents must coordinate while adapting to the evolving behavior of others. Evolutionary game theory can identify strategically stable outcomes, but stability under a population‑adjustment dynamic does not necessarily imply that finite‑sample learning agents can reach the same outcome through local reward feedback.

We illustrate the distinction with a transparent three‑agent governance game involving a government, a platform firm, and users. First, we derive the replicator dynamics for the fixed stage‑game incentives, expressed as

$$ \dot{x_i}=x_i\big[(A\mathbf{x})_i-\mathbf{x}^T A \mathbf{x}\big] $$

where $x_i$ denotes the proportion of strategy $i$ and $A$ is the payoff matrix. By evaluating the cooperative evolutionary basin on a symmetric initial‑condition grid, we obtain a volume $V_E=1.00$, indicating that cooperation is evolutionarily stable across the sampled space.

We then compare this with learning‑basin estimates for three decentralized value‑based learners under the same payoff environment: independent $\varepsilon$‑greedy Q‑learning ($\varepsilon$‑IQL), scaled Boltzmann exploration, and SA‑EA BQL. The learning basin is measured as the fraction of initial points that converge to the cooperative joint action on the identical grid. Empirically, $\varepsilon$‑IQL achieves a basin of $0.88$, while both scaled Boltzmann and SA‑EA BQL obtain $0.00$.

Diagnostic traces reveal that broader action diversity and non‑zero value separation can coexist with a failure to sustain cooperative joint action in this fixed configuration. These findings demonstrate that evolutionary stability and learning accessibility are distinct properties of a coupled game‑learning system.

The shared‑bike scenario serves as a motivating application; the broader contribution is a framework for comparing population‑level stability with the finite‑sample accessibility of cooperation under specified multi‑agent learning dynamics.

Review

Original Source: https://arxiv.org/abs/2609.27664

[h] Back to Home