NeFut Logo NeFut
Admin Login

[CS.AI] DrawingVQA: A Benchmark for Multi-Depth Visual-Textual Reasoning

Published at: 2026-07-20 22:00 Last updated: 2026-07-22 01:02
#AI #Machine Learning #Open Source

We introduce DrawingVQA, the first benchmark designed to evaluate multimodal large language models (MLLMs) on real-world construction drawings — a core medium in architecture, civil, and many other engineering practices.

Unlike natural images or schematic floor plans, construction drawings fuse abstract geometry, symbolic notation, tabular data, annotations, and domain-specific text, forming a uniquely complex visual-textual domain core to engineering workflows.

DrawingVQA bridges this gap with 33 'Issued for Construction' drawings and 92 expertly curated question-answer pairs, spanning three reasoning depths: perceptual understanding, contextual interpretation, and domain-expert reasoning.

To evaluate model capabilities, we present a dual categorization framework to jointly analyze performance across seven construction-engineering and four MLLM capability dimensions — the first to explicitly map engineering workflows to AI reasoning competencies.

Evaluations of state-of-the-art MLLMs reveal a substantial gap between model and expert performance, particularly at higher reasoning depths.

This benchmark lays a foundation for domain-specialized multimodal reasoning to allow for advancement on integration of AI-driven understanding and real-world engineering workflows.

Blogger's Review: The introduction of DrawingVQA marks a significant advancement in the capabilities of multimodal language models within the engineering domain, particularly in handling the complexities of construction drawings. This benchmark not only fills a gap in existing research but also provides a viable pathway for assessing and improving AI integration in real-world engineering applications.

Original Source: https://arxiv.org/abs/2607.15418

[h] Back to Home