This work investigates decentralized learning of socially optimal equilibria in finite normal‑form games played over time‑varying communication networks. Each agent only observes its own realized payoff, has no prior knowledge of the game, and can exchange low‑bandwidth messages with neighbors whose set changes over time.
We propose a networked decentralized optimal equilibrium learning dynamics. Agents generate randomized “satisfied/unsatisfied” signals from local payoff comparisons and share time‑stamped stacked tables instead of raw actions, payoffs, or local parameter estimates. Table fusion together with temporal majority reconstruction mitigates the burden of dynamic communication while preserving full decentralization.
We prove that, under utilitarian or proportional‑fair social welfare objectives and with an in‑phase exploration perturbation, the algorithm attains finite‑time logarithmic regret, i.e., $R(T)=O(\log T)$. Simulations demonstrate that the approach can effectively select socially desirable equilibria even when the communication network evolves.
Review