DrawingVQA: First AI Benchmark Tests Multimodal Models on Real-World Construction Drawings
Researchers introduced DrawingVQA, the first benchmark to evaluate multimodal large language models (MLLMs) on real-world construction drawings — a uniquely complex domain fusing abstract geometry, symbols, tables, and technical text. The benchmark uses 33 professional 'Issued for Construction' drawings and 92 expert-crafted questions to test AI's visual-textual reasoning in architecture and civil engineering.

Researchers have introduced DrawingVQA, the first benchmark designed to evaluate multimodal large language models (MLLMs) on real-world construction drawings — a core medium in architecture, civil engineering, and many other engineering disciplines. Unlike natural images or schematic floor plans, construction drawings fuse abstract geometry, symbolic notation, tabular data, annotations, and domain-specific text, forming a uniquely complex visual-textual domain that is central to engineering workflows.
DrawingVQA bridges this evaluation gap with 33 professional "Issued for Construction" drawings and 92 expert-crafted questions. The benchmark tests AI models on multi-depth visual-textual reasoning — from identifying simple symbols to interpreting complex relationships between geometry, annotations, and tabular data. This matters because better AI understanding of construction drawings can help architects and engineers quickly find specific details in large sets of drawings, reduce errors, and speed up project timelines.
The benchmark is available on arXiv, offering a resource for researchers working to improve AI's ability to interpret professional, real-world documents. While the technical details may be advanced, DrawingVQA represents a significant step toward making AI more useful in practical engineering and construction contexts.