NeFut Logo NeFut
Admin Login

[CS.AI] User-AI Mistreatment in Conversational Systems: Occurrence and Implications

Published at: 2026-09-15 22:00 Last updated: 2026-09-16 00:22
#AI #Machine Learning #LLM

Safety research often concentrates on harms generated by the model itself, yet users can also direct hostility, coercion, and adversarial pressure toward models. Accurately understanding this phenomenon is essential for interpreting model behavior, alignment drift, and real‑world deployment risks. In this work we audit 777,000 English LMSYS‑Chat‑1M conversations using two independent detectors: an eight‑category lexicon for hostility aimed at the model, and the dataset’s moderation signal. The two detectors capture weakly overlapping phenomena and are complementary. The lexicon flags insults, threats, and jailbreak coercion directed at the assistant, while moderation flags are dominated by requests for toxic content rather than model‑targeted hostility. Together they cover about 5% of user turns; after adjusting for measured precision the mistreatment rate toward the assistant is 0.90%. These absolute rates describe arena‑style evaluation traffic and should not be taken as deployment‑wide baselines. We observe a 13‑fold variation in user hostility across models, driven largely by the user base each model attracts rather than by model behavior: first‑turn hostility spreads far wider than post‑response hostility, and the extremes remain more than fifteenfold apart even after deduplicating opening prompts. Within conversations, assistant apologies consistently correlate with higher odds of next‑turn hostility under both detectors; this effect persists when restricting to non‑refused prior turns and jailbreak‑free conversations, and is positive in 20 of 23 models. Yet across models, more apologetic assistants receive less overall hostility. Temporally, coercive openings front‑load the first turn while affective hostility accumulates over a session. We release the lexicon, the detector cross‑validation pipeline, and all derived tables.

Review

Original Source: https://arxiv.org/abs/2609.13579

[h] Back to Home