This study examined whether large language models (LLMs) reproduce the motive attributions that accompany human moral judgments. Five mainstream LLMs (including GPT‑4, Claude, Llama‑2, etc.) and two human cohorts (N=125 and N=742) were asked to evaluate a physician who either stayed silent about fraudulent billing or reported it to a hospital, regulator, or newspaper. All models replicated the human ranking of the physician's moral character, but they systematically portrayed whistle‑blowers as more helpful, less self‑interested, and less hostile. In four of the five models, competitive motives were far less correlated with moral‑character judgments than in humans. Even when prompts reproduced the narrative context and demographic profiles of the human samples, model ratings changed little, indicating that agreement in average scores can mask differences in attributed motives, inter‑judgment relationships, and context sensitivity. Consequently, validating LLMs as simulated participants in psychological research requires testing psychologically informative response patterns, not merely average agreement.
Review