Large vision‑language models (VLMs) have made remarkable strides in multimodal understanding, yet their performance in educational contexts remains under‑examined. AI‑assisted language learning demands that models interpret artistic images, grasp their semantic, affective, and cultural meanings, and reason about visual context to enable meaningful interaction. Existing benchmarks mainly target real‑world photographs or domain‑specific educational reasoning, offering limited coverage of artistic educational content.
To bridge this gap, we introduce MUSE (Multimodal Understanding in Situated Education), a benchmark that decouples image annotation from question generation. This design allows diverse tasks with controllable difficulty while reducing annotation effort. MUSE comprises 12 tasks spanning visual perception, semantic interpretation, affective analysis, cultural understanding, and compositional reasoning. The curated image set deliberately emphasizes Singaporean and Southeast Asian multicultural contexts alongside Western art traditions, providing multiple themes and difficulty levels.
We evaluate a range of open‑source and proprietary models on MUSE. Results reveal substantial disparities across capability dimensions, especially in affective interpretation and compositional reasoning. Error analysis uncovers common failure modes such as misreading cultural symbols, under‑estimating emotional intensity, and broken cross‑modal reasoning chains. These findings highlight key challenges for building trustworthy multimodal models for education.
Review: MUSE offers a systematic, extensible platform for assessing multimodal understanding in educational settings, helping researchers pinpoint model weaknesses and drive progress toward more reliable AI assistants.