research

arXiv Paper Argues AI Leaderboards Structurally Exclude Global South Benchmarks Like IndicSUPERB and IrokoBench

Summarized by AI from reporting by ArXiv cs.AI, published under our editorial policy.

A new position paper on arXiv argues that AI leaderboards are structurally ill-suited to serving the Global South due to a lack of independent governance, conflict-of-interest policies, and metric evolution mechanisms. The paper emphasizes that the barrier is not missing data—high-quality regional benchmarks like IndicSUPERB, MILU, and LAHAJA for India, IrokoBench for Africa, and AlGhafa for Arabic already exist—but rather institutional design and commercial pressure.

A diverse group of people using AI technology on various devices.

Key takeaways

  • AI leaderboards lack independent governance and conflict-of-interest policies, leading to the structural exclusion of Global South benchmarks.
  • High-quality regional benchmarks like IndicSUPERB, MILU, LAHAJA, IrokoBench, and AlGhafa already exist but are not included in global leaderboards.
  • The paper argues the barrier is institutional design and commercial pressure, not missing data.
  • The exclusion of regional benchmarks can result in AI models that perform poorly in the Global South, exacerbating existing inequalities.

A new position paper on arXiv (2608.18117) argues that AI leaderboards are structurally ill-suited to serving the Global South. The paper highlights that these leaderboards lack independent governance, conflict-of-interest policies, and mechanisms for updating their metrics. The authors state that the barrier is not a lack of data, as high-quality regional benchmarks already exist: IndicSUPERB, MILU, and LAHAJA for India; IrokoBench for Africa; and AlGhafa for Arabic. The problem, they argue, is institutional design: global leaderboards do not include these benchmarks, and no governance mechanism compels them to do so. Commercial pressure corrects leaderboards toward Western-centric metrics, further entrenching the exclusion.

Why Global Leaderboards Exclude Regional Benchmarks

AI leaderboards are widely used to evaluate and compare the performance of AI models. However, these leaderboards are predominantly focused on Western-centric benchmarks, which do not adequately represent the linguistic and cultural diversity of the Global South. For instance, benchmarks like IndicSUPERB, which is designed to evaluate AI performance on Indian languages, are not included in major global leaderboards. The paper argues that this exclusion is not accidental but stems from the lack of independent governance and conflict-of-interest policies in the organizations that run these leaderboards. Without governance mechanisms, commercial pressures from large AI companies in the Global North dictate which benchmarks are prioritized.

Real-World Impact on Users in the Global South

The exclusion of regional benchmarks has real-world implications. AI models that are not tested on diverse linguistic and cultural data may perform poorly when deployed in the Global South. For example, a language model trained primarily on English data may struggle to understand and generate text in Indian languages, leading to poor user experiences. This disparity can exacerbate existing inequalities and limit the benefits of AI for people in these regions. The paper notes that the problem is not that regional benchmarks don't exist—they do—but that the institutional structure of global leaderboards prevents their adoption.

Proposed Solutions: Independent Governance and Metric Evolution

To address this issue, the paper calls for the establishment of independent governance bodies for AI leaderboards. These bodies should have the authority to include regional benchmarks and ensure that leaderboards evolve to reflect the diverse needs of the global population. Additionally, conflict-of-interest policies should be implemented to prevent commercial pressures from influencing the selection of benchmarks. The paper also advocates for mechanisms for metric evolution, allowing leaderboards to update their evaluation criteria over time as new regional benchmarks emerge. Users can also advocate for greater inclusion by supporting organizations that develop and promote regional benchmarks.

For those interested in learning more about this issue, the full paper is available on arXiv under identifier 2608.18117.

Frequently asked

What specific regional benchmarks does the paper mention as already existing?
The paper mentions IndicSUPERB, MILU, and LAHAJA for India; IrokoBench for Africa; and AlGhafa for Arabic.
Does the paper say the problem is a lack of data from the Global South?
No, the paper explicitly states that the barrier is not missing data—high-quality regional benchmarks already exist—but rather institutional design and governance failures.
What does the paper propose as a solution to this problem?
The paper calls for the establishment of independent governance bodies for AI leaderboards, conflict-of-interest policies, and mechanisms for metric evolution to ensure regional benchmarks are included.