Mentalization refers to the ability to infer others' beliefs and intentions to guide one's own choices, a core cognitive function underlying human social interaction. Recent large language models (LLMs) exhibit behavior on theory‑of‑mind tasks that resembles humans, yet it remains unclear whether they can use mentalization to drive adaptive behavior.
This study employs two economic games together with cognitive computational modeling to uncover the latent strategies behind LLM mentalization. Experiments tested 2,099 LLM agents across four model families (DeepSeek, GPT‑4.1, GPT‑5, Gemini 2.0 Flash) against opponents of varying sophistication, and examined whether a prompting strategy designed to elicit strategic reasoning improves performance. Results were benchmarked against data from 251 human participants.
Across both games, LLMs displayed clear behavioral and computational signatures of mentalizing, but their performance varied markedly by model provider and scale. Strategic prompting generally enhanced performance by inducing higher‑order reasoning, though the magnitude of improvement differed between tasks. Notably, GPT‑5 agents flexibly adjusted their recursive reasoning depth in response to increasingly sophisticated opponents, ultimately outperforming human participants.
The research demonstrates significant differences in mentalization capacities among LLMs and highlights cognitive computational modeling as a formal method for comparing intelligence between humans and machines.
Blogger's Review: By combining rigorous experimental design with detailed model analysis, this work provides some of the first systematic evidence that state‑of‑the‑art LLMs possess adjustable mentalizing reasoning, opening promising avenues for social interaction and human‑machine collaboration.