AI systems are deployed worldwide, yet they often fail to meet the needs of culturally diverse users. Prior work has mainly evaluated cultural understanding in text‑only settings or by recognizing isolated visual artifacts such as food or clothing, leaving visual norm understanding—reasoning about observable behaviors through local social norms—largely unexplored. To address this gap we introduce NormViz‑Bench, a human‑validated benchmark comprising 3,268 contrastive image pairs (6,536 images) from 16 countries. Each pair differs only in culturally relevant behavior elements—objects, attributes, spatial relations, and actions—and every image is labeled as conforming, violating, or irrelevant to local norms. Pair‑level evaluation requires both images to be correctly classified, preventing reliance on superficial visual shortcuts. Even the strongest vision‑language models, Gemini 3.0 Flash and Qwen2.5 VL 7B, succeed on only 26.6% and 21.6% of pairs, respectively, struggling most with identifying violating and culturally benign behaviors. To bridge this gap we release NormViz‑Train, a training set of 64k images paired with explanatory texts. Although absolute performance remains low, the data provide a promising avenue for improving visual norm understanding.
Review