For experienced operators, reading a gauge is almost trivial: it requires little domain knowledge, imposes low cognitive load, and yields highly repeatable results. In contrast, multimodal large language models (MLLMs) remain unreliable for continuous‑valued measurement despite strong performance on generic multimodal benchmarks. Existing benchmarks isolate measurement from realistic, knowledge‑grounded contexts, lacking specialized instruments, real‑world noise, and diagnostic annotations, which limits realism and hampers root‑cause analysis.
We introduce InSituMeasure, a benchmark designed to evaluate situated measurement grounding. It comprises 2,922 real industrial monitoring images spanning eight functional categories of professional engineering instruments. Each image is densely annotated with gauge attributes and noise tags to support failure diagnosis. We define three metric families: (1) numerical accuracy within predefined tolerances and unit consistency; (2) rejection of fake or unanswerable tasks; (3) alignment between model failures and annotated error factors.
Across 24 state‑of‑the‑art MLLMs, the best joint value‑unit accuracy reaches only 25.7% and the confidence‑diagnosis F1 scores 51.8%, highlighting a substantial gap between general multimodal competence and reliable situated measurement. Error analysis reveals three dominant failure modes: text‑induced shortcuts, overconfident responses, and authentic industrial noise—including mixed disturbances, viewpoint deviation, occlusion, and environmental interference.
Review