We examined cooperative tendencies of frontier generative AI (Gen AI) systems using the iterated prisoner's dilemma, manipulating counterpart reputation (positive, unknown, negative), strategy (extortion versus generosity), and non‑verbal emotional cues (facial expressions indicating competitive or cooperative appraisals).
In the first study we evaluated non‑reasoning models Claude 3.5, Gemini 2.0 Flash, and GPT‑4o. Cooperation was systematically shaped by all three factors, reproducing patterns long documented in human behavioral research, yet each model weighted the factors differently.
The second study focused on reasoning models Claude 4.6, Gemini 3, and GPT‑5.2. These models relied more heavily on strategy and reputation, the Potemkin effect observed in non‑reasoning models nearly vanished (diagnostic harmony game yielded almost uniform cooperation), and emotion played a conditional role consistent with hierarchical cue‑weighting rather than a simple loss of social sensitivity. Reasoning models also displayed heterogeneous end‑game behavior, ranging from sustained cooperation to systematic last‑round defection, revealing model‑specific exploitability profiles relevant for negotiation and other multi‑round interactions.
Together the findings characterize generative AI as increasingly sophisticated yet heterogeneous social actors and underscore the practical value of standardized cooperation benchmarks for responsible deployment in interactive, socially consequential settings.
Review