NeFut Logo NeFut
中 Admin Login

[CS.AI] Robust Nash Alignment under Preference Uncertainty

Published at: 2026-10-02 22:00 Last updated: 2026-10-06 12:11
#algorithm #AI #LLM

Preference‑based alignment methods usually optimize against a single preference model, which makes them fragile when pairwise preferences are noisy, heterogeneous, or shift after deployment. To tackle these issues we introduce Robust Nash Alignment, a game‑theoretic framework for aligning with uncertain pairwise preferences. The major learner seeks a policy that maximizes the worst‑case win rate against an adversarial competitor and any preference kernel within an ambiguity set around a nominal preference. When the ambiguity set captures preference uncertainty, the robust objective directly yields a certified lower bound on worst‑case performance. Optimizing this objective is computationally demanding. We therefore design a four‑player primal‑dual proxy game involving a leader policy, a follower policy, an adversarial kernel, and a dual variable, and solve it with a single‑loop optimistic mirror descent‑ascent algorithm. We show the proxy always lower‑bounds the truncated hard‑constrained objective, quantify the proxy‑to‑hard gap, and identify an exactness condition under which the proxy recovers the robust objective. Moreover we prove an average‑iteration convergence of the proxy‑game duality gap at $\mathcal{O}(1/\sqrt{T})$, implying a near‑optimal robust policy for the original objective. Experiments on controlled tabular games and LLM alignment with uncertain preferences validate the convergence theory and demonstrate improved performance over nominal baselines.

Review

Original Source: https://arxiv.org/abs/2610.00715

[h] Back to Home