#benchmark

Benchmark

116 stories tagged Benchmark

A diverse group of people using AI technology on various devices.
research

arXiv Paper Argues AI Leaderboards Structurally Exclude Global South Benchmarks Like IndicSUPERB and IrokoBench

A new position paper on arXiv argues that AI leaderboards are structurally ill-suited to serving the Global South due to a lack of independent governance, conflict-of-interest policies, and metric evolution mechanisms. The paper emphasizes that the barrier is not missing data—high-quality regional benchmarks like IndicSUPERB, MILU, and LAHAJA for India, IrokoBench for Africa, and AlGhafa for Arabic already exist—but rather institutional design and commercial pressure.

via ArXiv cs.AI#AI#Benchmark#Research
A graph comparing FLOPs to actual execution times for different AI operations.
research

FLOPs vs Real Work: Why AI Efficiency Metrics Need Replication Studies

A new ArXiv study challenges the common use of FLOPs (Floating Point Operations) to measure AI efficiency, finding that operations with the same FLOPs can have execution times varying by up to 30% due to differences in parallelization. The researchers replicated a previous study to demonstrate why FLOPs alone are a misleading metric, calling for more nuanced benchmarks that capture real-world performance.

via ArXiv cs.AI#Benchmark#Research
DrawingVQA: First AI Benchmark Tests Multimodal Models on Real-World Construction Drawings
research

DrawingVQA: First AI Benchmark Tests Multimodal Models on Real-World Construction Drawings

Researchers introduced DrawingVQA, the first benchmark to evaluate multimodal large language models (MLLMs) on real-world construction drawings — a uniquely complex domain fusing abstract geometry, symbols, tables, and technical text. The benchmark uses 33 professional 'Issued for Construction' drawings and 92 expert-crafted questions to test AI's visual-textual reasoning in architecture and civil engineering.