A pool of language models can collaborate and improve collectively by learning from each other's responses. The interactions depend on the instructions used during training. Existing approaches typically sample instructions uniformly, even though an instruction's usefulness changes as models get better: an instruction that once caused quality differences may later be answered equally well, while a previously hard instruction may start providing a useful learning signal.
We propose Stackelberg Alignment, a game‑theory‑inspired leader‑follower framework that turns instruction selection into an adaptive curriculum. An EXP3 bandit acts as the leader, allocating a fixed sampling budget across instructions and updating its sampling distribution with a reward that combines instruction difficulty and response discriminability. The language models act as followers: they respond to the selected instructions, evaluate each other's replies, and learn from the resulting preference signals via DPO or GRPO. The framework uses Elo‑style reputation‑weighted peer judgment and reputation‑based opponent matching to ensure reliable and competitive interactions.
Experiments across three heterogeneous model pools and twelve benchmarks covering scientific discovery, reasoning, code, instruction following, and knowledge show that Stackelberg Alignment achieves the highest macro‑average across the pools, outperforming the strongest training‑time baseline by up to 7.4% and the best static inference baseline by 12%‑25%. Analysis confirms that the adaptive leader concentrates duels on the most informative instructions, and ablations demonstrate that both reputation‑weighted judgment and reputation‑based matching improve the effectiveness of multi‑LLM evolution.
Review: By integrating game theory with an adaptive curriculum, this work offers a systematic solution for multi‑model collaborative learning, and the extensive experiments validate its robustness and advantage across diverse tasks and model configurations.