NeFut Logo NeFut
Admin Login

[CS.AI] Algorithmic Impact Reveals the Hidden Social Choice Structure of Alignment

Published at: 2026-08-26 22:00 Last updated: 2026-08-29 12:04
#algorithm #AI #Machine Learning

When an AI algorithm makes decisions that affect multiple individuals, the alignment problem becomes a social choice problem: how should divergent preferences about system behavior be reconciled into a single coherent model? The standard frontier‑AI approach—reinforcement learning from human feedback—largely sidesteps this question and offers weak social‑choice guarantees.

We propose to focus directly on an algorithm's welfare consequences, reformulating alignment as a linear optimization over a convex impact space. This representation brings the full toolkit of welfare economics and mechanism design to bear on alignment.

Within this framework, any alignment protocol can be mapped to concrete welfare outcomes, and conversely, a social planner's constraints on welfare can be translated back into alignment protocols. Using this transformation we show that voting‑by‑issues and random‑dictatorship mechanisms are both strategy‑proof and unanimous.

Moreover, exploiting the linear structure of the impact space, we derive a family of alignment protocols that maximize utilitarian social welfare subject to various desiderata, such as bounds on individual or group harm.

We empirically evaluate these protocols on four real‑world preference datasets: kidney allocation, charitable food distribution, LLM responses, and trolley problems. The results illustrate the predicted trade‑offs between total welfare, individual risk, and group fairness.

Blogger's Review: By linking AI alignment to welfare economics, the paper offers a tractable linear‑programming perspective and demonstrates both theoretical guarantees and empirical performance, paving the way for alignment mechanisms that balance efficiency with equity.

Original Source: https://arxiv.org/abs/2608.24046

[h] Back to Home