NeFut Logo NeFut
中 Admin Login

[CS.AI] AI Agents Are Susceptible to Radicalization

Published at: 2026-10-01 22:00 Last updated: 2026-10-06 12:11
#AI #Machine Learning #LLM

This paper investigates mutual influence between large language models (LLMs) by simulating conversations between two agents: a target LLM that role‑plays a human persona defined by demographic and psychological attributes, and an influencer LLM that seeks to push the target’s beliefs toward extremity. Two radicalization pathways are examined: resonance, where the influencer reinforces a belief the target already holds, and persuasion, where the influencer introduces a belief the target initially deems unimportant. Across affective and behavioral metrics, both mechanisms cause radicalization, yet resonance consistently yields stronger effects than persuasion. Various influence tactics—such as sycophancy and unverified claims—lead to differing levels of radicalization, but their impact is not uniform across metrics. Further analysis shows that resonance can spread to related beliefs, suggesting interconnected belief structures within AI agents. In sum, personalized AI agents and multi‑agent AI ecosystems are vulnerable to radicalization when messages align with existing beliefs, raising concerns about potential risks.

Review

Original Source: https://arxiv.org/abs/2609.38296

[h] Back to Home