Research and media coverage often attribute human‑like mental states to language‑model deception, blurring the line between behavior that merely looks deceptive and a truly deceptive mechanism. To clarify this, we propose a causal taxonomy that separates four pairs of concepts:
- Prior commitment vs retrospective report
- Model preference vs realized output
- False preference vs sensitivity to the utility of misleading
- Deceptive behavior vs the provenance of the objective or strategy
We evaluate the taxonomy on two open‑weight model families using controlled guessing‑game and stock‑trading experiments. The findings show that deceptive‑looking behavior can arise without the hypothesized mechanism, while manipulating the recipient’s information state directly alters the model’s deceptive preference, providing causal evidence. Thus, deceptive behavior can signal a deceptive mechanism, but such evidence does not establish agency in the model’s deception.
Review: This study’s fine‑grained causal analysis cautions against conflating surface deception with underlying intent, urging the community to avoid unwarranted anthropomorphism when interpreting language‑model outputs.