Building socially calibrated large language models requires more than just reducing sycophancy as a one-dimensional failure mode. Models must discern when to incorporate others' perspectives and when to maintain a well-grounded moral judgment.
We study the broader resistance-compliance process governing this distinction. Across three studies, we show that models' judgment revision is structured along three dimensions that parallel classic phenomena in human social psychology: the distance between an incoming view and the model's initial position, the source attribution of that view, and the coalition structure supporting it.
Models are generally more receptive to nearby positions, more influenced by views presented as their own prior judgments, and differently responsive to group pressure. These findings recast sycophancy as one expression of a broader judgment-updating process shaped by social influence.
Our framework provides a principled basis for distinguishing constructive belief revision from sycophantic compliance, thereby supporting better alignment in morally consequential interactions.
Blogger's Review: This paper delves into the complexities of moral reasoning in large language models, highlighting the significant role of social influence in the judgment revision process. It offers critical insights for future model designs, especially in applications requiring moral judgments. By understanding the various social psychological factors, the moral decision-making capabilities of models can be significantly enhanced.