Previous AI alignment research has largely targeted first‑order social norms—teaching models what actions are acceptable or prohibited (e.g., “do not steal”). True social intelligence, however, also requires anticipating who will enforce a norm and how (e.g., public shaming or imprisonment). These second‑order expectations are called metanorms and dictate people’s responses when rules are broken.
We introduce a framework for evaluating metanorm reasoning in large language models (LLMs) along two axes: emotional appraisal and behavioral response. Two classification tasks are defined: predicting self‑regulation by violators and other‑regulation by observers. To support this, we release the multi‑perspective dataset NormReact, containing 450 norm‑violation scenarios annotated for emotions and behavioral responses, with annotations split by violator gender and observer social closeness.
Across six representative LLMs, we find a systematic bias toward over‑predicting negative sanctions—models often forecast punishment even in cases where humans would expect inaction. Alignment with human judgments deteriorates as social distance between observer and violator increases.
These results suggest that AI systems deployed in norm‑sensitive domains such as conflict mediation or policy simulation may present a distorted picture of social regulation: punishment is over‑represented while tolerance, restraint, and relational calibration that characterize real‑world norm enforcement are under‑represented.
Review