NeFut Logo NeFut
Admin Login

[CS.AI] From Deceptive Outputs to Deceptive Mechanisms: A Causal Framework for Language-Model Deception Research

Published at: 2026-09-05 22:00 Last updated: 2026-09-06 01:02
#algorithm #AI #Machine Learning

Research and media coverage often attribute human‑like mental states to language‑model deception, blurring the line between behavior that merely looks deceptive and a truly deceptive mechanism. To clarify this, we propose a causal taxonomy that separates four pairs of concepts:

We evaluate the taxonomy on two open‑weight model families using controlled guessing‑game and stock‑trading experiments. The findings show that deceptive‑looking behavior can arise without the hypothesized mechanism, while manipulating the recipient’s information state directly alters the model’s deceptive preference, providing causal evidence. Thus, deceptive behavior can signal a deceptive mechanism, but such evidence does not establish agency in the model’s deception.

Review: This study’s fine‑grained causal analysis cautions against conflating surface deception with underlying intent, urging the community to avoid unwarranted anthropomorphism when interpreting language‑model outputs.

Original Source: https://arxiv.org/abs/2609.04166

[h] Back to Home