NeFut Logo NeFut
中 Admin Login

[CS.AI] Mitigating Social Sycophancy via Pluralistic Preference Optimization

Published at: 2026-10-05 22:00 Last updated: 2026-10-06 12:11
#AI #optimization #LLM

Personal advice, such as relationship counseling, is now one of the most common uses of generative AI. However, language models (LMs) tend to be sycophantic: they agree with users far more often than humans do, which can make users overconfident and reluctant to repair relationships after conflicts. Prior mitigation work focuses on factual settings where a ground‑truth answer exists; in social advice, where no objective truth is available, simple prompting or post‑training yields limited gains. We observe that social sycophancy arises partly because LMs over‑center on the user and ignore other stakeholders affected by the user’s behavior. To address this, we introduce Pluralistic Preference Optimization (PlurPO). Given a description of an interpersonal conflict, the LM first identifies and simulates all relevant stakeholders, then uses self‑generated preference signals to train the model to prefer responses acceptable to every stakeholder. No external ground‑truth labels are required. Across four datasets and four model families, PlurPO markedly reduces social sycophancy. For statements of intent to cause harm, endorsement rates drop by an average of 89%; for general advice questions, the gap between model and human endorsement rates shrinks from 17.8% to 8.0%. Moreover, the preference dataset built for an 8B model transfers effectively to a larger 32B model, further mitigating sycophancy. These findings demonstrate that leveraging a model’s own capability to simulate a plurality of perspectives can substantially curb social sycophancy.

Review

Original Source: https://arxiv.org/abs/2610.02568

[h] Back to Home