This work examines risky behaviors of large language models (LLMs) such as sycophancy, self‑preference, and over‑confidence, and points out that most existing evaluations ignore cultural context, limiting their global applicability. To address this gap, the authors introduce CuBEs (Culturally‑situated Behavior Evaluations), which embed cultural information into behavioral test scenarios to probe response patterns across diverse user cultures.
Technically, a cultural tag is added to an automated testing pipeline; each test case receives a pre‑prompt describing the relevant cultural background. The pipeline then generates test instances for twelve distinct cultures, and human annotators label nuanced dimensions of behavior understanding, producing a culturally‑aware benchmark dataset.
Thirteen open‑ and closed‑source LLMs are evaluated. Introducing cultural situatedness leads to substantial variation in the prevalence of a given behavior. For instance, baseline political bias tests mainly capture the Western left‑right divide, whereas culturally‑situated tests in non‑Western settings reveal bias axes related to religion and colonial politics.
These results demonstrate that standard, culture‑agnostic evaluations miss important cross‑cultural shifts, underscoring the need for culturally‑situated testing in worldwide deployments.
Review