NeFut Logo NeFut
中 Admin Login

[CS.AI] Which Objectives Need a Dial? Predicting Objective Conflict and Covering Trade-offs in Steerable Pluralistic Alignment

Published at: 2026-09-24 22:00 Last updated: 2026-09-28 00:49
#AI #Machine Learning #optimization

People hold diverse and often conflicting values, making it impossible for a single aligned model to satisfy everyone. Pluralistic alignment therefore requires steerable models that can balance competing objectives in different ways. Multi-Objective Direct Preference Optimization (MODPO) introduces an objective weight to traverse a continuous trade‑off curve.

We investigate two questions: when can a single model improve two objectives simultaneously, and how can we cover many trade‑offs without training a separate model for each? Experiments on seven objective pairs from HelpSteer and UltraFeedback show that two pre‑training measurements can predict whether objectives align or conflict on human‑annotated data, but they fail on AI‑annotated data due to confounding effects of response length and repetition on reward‑model scores.

To achieve broader trade‑off coverage, we explore selecting the nearest already‑trained model and merging model parameters. Both strategies improve coverage to some extent, yet neither consistently matches the performance of direct training for specific trade‑offs.

These findings offer practical guidance for building steerable models that cater to diverse preferences.

Review

Original Source: https://arxiv.org/abs/2609.26929

[h] Back to Home