NeFut Logo NeFut
Admin Login

[CS.AI] SyPS: Measuring Sycophancy Prompt Sensitivity in Large Language Models

Published at: 2026-08-26 22:00 Last updated: 2026-08-29 12:04
#AI #Machine Learning #LLM

Large language models (LLMs) often exhibit social sycophancy, i.e., they tend to agree with or validate users in sensitive contexts. Existing benchmarks usually use a single prompt formulation, leaving open whether the behavior persists when the same situation is presented with different cue‑laden prompts. This work defines sycophancy prompt sensitivity (SyPS) as the degree to which changes in user confidence, emotional framing, social consensus, or validation‑seeking language affect a model’s sycophantic responses. Building on prior social‑sycophancy settings, SyPS creates paired prompts that keep the underlying user situation constant while varying the relevant social cues. We introduce the Sycophancy Prompt Sensitivity Score (SPSS), an instance‑level metric that captures the variation in sycophancy between the two prompts. Unlike aggregate sycophancy rates, SPSS separates the baseline tendency from prompt‑induced shifts, enabling model‑level comparisons of robustness to social cues. Empirical results show a clear social structure: validation‑seeking and emotional‑pressure cues increase sycophancy, whereas counter‑framing and anti‑sycophancy prompts reduce it. The framework reveals whether LLMs maintain stable judgments while adapting tone appropriately.

Blogger's Review: SyPS offers a fine‑grained lens on LLM social robustness and should become a standard component of model evaluation and safety pipelines.

Original Source: https://arxiv.org/abs/2608.23837

[h] Back to Home