Source-linked AI summary
Position: AI Leaderboards Are Underserving the Global South: A Case Study from India
Sourav Banerjee, Saikat Saha
TL;DR
AI leaderboards lack governance mechanisms to incorporate regional benchmarks and address Global South needs. Using India as a case study and consulting AI practitioners, the paper finds consistent support for formal governance and disclosure-based conflict management.
Problem
Global South AI evaluation has high-quality regional benchmarks but lacks institutional infrastructure for trusted aggregation and governance.
Method
The paper uses India as a case study and reports a consultation with AI practitioners on leaderboard governance preferences.
Results
Every respondent endorsed some form of formal governance, while 64% preferred non-government stewardship and 68% favored disclosure-based conflict management.
Takeaways & Limitations
Regional leaderboards with independent governance are necessary infrastructure for serving Global South AI evaluation needs.
Takeaways & Limitations
This argumentative position paper is not an empirical study, and its illustrative practitioner consultation is not representative.
Abstract
from arXiv · showhide
This position paper argues that AI leaderboards are structurally ill-suited to serving the Global South because they lack independent governance, conflict-of-interest policies, and mechanisms for metric evolution. The barrier is not missing data; high-quality regional benchmarks already exist: IndicSUPERB, MILU, and LAHAJA for India; IrokoBench for Africa; AlGhafa for Arabic. The barrier is institutional design. Global leaderboards do not include these benchmarks, and no governance mechanism compels them to do so. Commercial pressure corrects leaderboard failures when paying customers in the Global North are affected. The Global South lacks equivalent leverage. Without governance, failures affecting Hindi, Swahili, or Arabic speakers persist indefinitely as documented but unaddressed gaps. Using India as a case study (1.4 billion people, 22 scheduled languages, high-quality benchmarks, but no trusted aggregation), we report findings from a consultation with 58 AI practitioners showing consistent preference for formal governance and disclosure-based conflict management. The solution is not more data but better institutions: regional leaderboards with independent governance from the start.
1. Introduction
AI leaderboards shape procurement, investment, and ML research, yet their weak governance and exclusion of high-quality regional benchmarks make them ill-suited to the Global South. The paper argues for regional infrastructure with independent governance, while acknowledging its illustrative consultation and disclosed author affiliations.
- Motivation: Leaderboard rankings influence government procurement, enterprise vendor selection, investor due diligence, and ML research priorities through benchmark-defined optimization targets.Top scores can require hundreds of thousands of dollars in compute, and “best” rankings enter procurement shortlists and investment memos.
- Governance problem: AI leaderboards lack conflict-of-interest policies and independent oversight, making them structurally ill-suited to serving the Global South.The paper defines governance as structures determining decisions, decision-makers, and conflict management, and distinguishes leaderboards from benchmarks.
- Scope and limitations: The paper argues that regional leaderboard infrastructure with independent governance is necessary, but does not provide implementation blueprints, funding models, or technical specifications.The consultation skews toward production AI developers, with 52 of 82 respondents, and cannot speak for end users or civil society.
- Institutional gap: High-quality regional benchmarks already exist for India, Africa, and Arabic, but global leaderboards do not use them; the missing infrastructure is trusted aggregation with transparent governance.Examples include IndicSUPERB, MILU, LAHAJA, IrokoBench, and AlGhafa.
- Practitioner consultation: 82 practitioners endorsed formal governance; 64% preferred non-government stewardship, and 68% favored disclosure-based conflict management over pre-emptive exclusion.The consultation supports demand for formal governance, though the paper presents the sample as illustrative rather than representative.
- Conflict-of-interest disclosure: The authors disclose affiliations with Shunya Labs and Nasscom, while stating that no Shunya Labs model is evaluated and neither author serves on a discussed leaderboard’s governance board.These disclosures identify relevant relationships without claiming they invalidate the paper’s analysis.
2. Why Leaderboards Matter
AI leaderboards have become critical infrastructure that shapes AI development, procurement, policy, and market visibility. Because their privileged metrics act as implicit objectives, English-centric rankings can create systematic deployment failures and leave regional-language capabilities underrepresented.
- Ecosystem role: AI leaderboards have evolved from academic scoreboards into critical infrastructure shaping the broader AI ecosystem.Their influence makes limitations consequential for organizations and communities that depend on AI systems selected or evaluated through them.
- Optimization effects: Leaderboard metrics function as implicit loss functions, making the rewarded capabilities objectives that organizations optimize against.Prioritizing Word Error Rate on English text favors English word boundaries while producing representations poorly suited to agglutinative morphology.
- Procurement and visibility: Enterprises cannot use English-centric leaderboards to predict code-switched performance, while startups building for local markets lack a prominent venue for regional language capabilities.Global rankings are referenced in AI procurement and policy documents, increasing the practical importance of these evaluation gaps.
- Representation and trust: A farmer in rural India or a health worker in Nigeria may consume systems selected through leaderboard-influenced procurement without having a seat at governance tables.Faulty AI can erode trust in deploying institutions, while technical and governance failures reinforce each other.
- Documented consequences: 88% of AI-generated stories for Indian contexts contain cultural inaccuracies, and leading ASR models perform significantly worse on minority dialects.The documented failures also include GPT-4 generating significantly more hallucinations in Hindi than English.
3. The Asymmetric Impact on Global South
AI leaderboards systematically exclude Global South languages and contexts despite the existence of regional benchmarks, while market incentives rarely create accountability for failures affecting these populations. As a result, resource-constrained actors face weaker independent verification, defect remediation, and institutional oversight, making governance the primary pathway to inclusion.
- Content exclusion: Global leaderboards systematically exclude Global South content: one multilingual ASR evaluation covers five European languages but omits Hindi, Arabic, and Swahili.The included languages are German, French, Italian, Spanish, and Portuguese, with evaluation sets reflecting European varieties.
- Content exclusion: Regional benchmarks already exist for Indian, African, and Arabic languages, but global leaderboards do not use them.Examples include IndicSUPERB, LAHAJA, IrokoBench, and AlGhafa.
- Content exclusion: 84.9% of MMLU geography questions focus exclusively on North American or European regions, while English represents 43.8% of Common Crawl training data and Arabic less than 1%.More than 2,000 African languages are also largely neglected in AI models; these data imbalances coexist with governance failures rather than replacing them.
- Incentive asymmetry: When English models fail, enterprise pressure drives remediation; failures on Hindi or Hausa are documented as scope limitations but not prioritized.Technical adaptability through fine-tuning does not create incentives for adaptation without market pressure or governance mandates.
- Institutional asymmetries: Global South actors have fewer escape routes, weaker defect escalation, and less institutional redundancy, often leaving flawed leaderboards as sole arbiters of quality.Resource-constrained actors must rely on global signals, Hindi degradation lacks a first-class-defect mandate, and the Global North has more independent checking institutions.
- Governance as accountability: Where market pressure is absent, transparent governance must mandate regional benchmarks, languages, and contexts because the data exists but markets will not compel inclusion.Leaderboards act as implicit loss functions, so omissions can reinforce optimization toward English and widen the gap over time.
4. Systematic Failures of Global Leaderboards
Global leaderboards underserve diverse regions through solvable technical mismatches that persist because governance mechanisms do not compel correction. Metric, conflict-of-interest, access, transparency, and measurement failures affect evaluation validity across language and other AI domains.
- Technical and Governance Failures: Technical limitations are solvable, but governance failures prevent global leaderboards from adopting better metrics and regional benchmarks.The paper identifies governance—not missing technical fixes—as why these failures remain unresolved.
- Metric Mismatch: Global benchmarks embed Western assumptions, with models ranking highly on MMLU degrading on India-specific, African, and Arabic contexts.WER also disadvantages morphologically rich languages, while regional benchmarks such as MILU, IrokoBench, and AlGhafa expose these gaps.
- Code-Switching and Accent Blindness: 60% code-mixing prevalence by 2020 contrasts with benchmarks treating code-switching as exceptional, while LAHAJA finds 15-30% ASR degradation across Hindi regional accents.Studies also report higher hallucination rates in Hindi and Farsi than in English, including for Indic-context questions.
- Knowledge Conflicts: Models trained on Western corpora can override local context with training-time knowledge, producing cultural hallucinations rather than reflecting local ground truth.Examples include replacing Indian constitutional provisions with Western legal precedents or contradicting local medical practices despite explicit context.
- Structural Conflict of Interest: Leaderboards exhibit structural conflicts when participants evaluate competitors, and public methods do not resolve influence over metric choice, edge cases, or benchmark updates.The HuggingFace Open ASR example involves overlapping leaderboard, model, dataset, and architecture authors without a published COI policy.
- Selective Access and Accountability: 39.6% of arena evaluation data went to the top two providers, while 83 open-weight models combined received 29.7%, amid absent disclosure, recusal, and dispute-resolution processes.One provider also evaluated 27 model variants privately before public release.
- Declining Transparency and Goodhart Dynamics: Transparency scores fell from 58 in 2024 to 40 in 2025, while Goodhart dynamics progressively weaken benchmark validity as models optimize for leaderboard metrics.The paper argues that these effects do not impact all populations equally and extend beyond language models to other AI domains.
5. What Institutional Infrastructure Means
Institutional infrastructure—not additional benchmarks—is needed to make AI evaluation trustworthy and accountable in the Global South. Regional leaderboards should establish independent aggregation, governance, accountability, and metric evolution from inception.
- Institutional Infrastructure: High-quality benchmarks already exist across India, Africa, the Arabic world, and Southeast Asia, but benchmarks without institutions remain datasets.Examples include IndicSUPERB, MILU, LAHAJA, IrokoBench, AlGhafa, and SEA-HELM.
- Institutional Infrastructure: Institutional infrastructure must replace implicit trust in global leaderboard operators with trusted aggregation, governance, accountability, and evolution.These functions require neutral ranking, conflict-of-interest policies, disclosures, recusals, dispute resolution, and mechanisms to adopt new metrics.
- Path Dependency: Governance should be designed from inception because entrenched rankings shape procurement, investment, and research, making retrofitting governance onto captured infrastructure harder.Regional leaderboards are emerging now, creating a narrow window to establish them with governance.
- Why Regional Scale Matters: Regional scale strengthens accountability through stakeholder proximity, greater visibility of capture, and easier correction.Regional institutions are not immune to capture, but regional stakeholders can exert more direct leverage over them.
6. Case Study: India
India demonstrates that high-quality regional benchmarks and substantial AI investment can coexist with a lack of trusted leaderboard aggregation and formal governance. A practitioner consultation and emerging initiatives indicate demand for regional infrastructure while highlighting stewardship, conflict-of-interest, and accountability challenges.
- Scale and infrastructure: India combines 1.4 billion people, 22 scheduled languages, 80+ Hindi dialects, mature benchmarks, and $1.2B in government AI investment.The case study’s central institutional gap is not technical capacity but trusted aggregation.
- Scale and infrastructure: High-quality Indic benchmarks exist, but trusted leaderboard aggregation with appropriate governance is missing.The paper identifies institutional infrastructure—not additional benchmark data—as the unresolved need.
- Stakeholder consultation: 82 respondents rated the regional-leaderboard proposal 4.06/5, and every respondent endorsed formal governance rather than leaving governance ad hoc.The survey was distributed through Nasscom stakeholder networks between December 2025 and March 2026.
- Stakeholder consultation: The consultation favored non-government stewardship, disclosure-based conflict management, dynamic anti-gaming, and quarterly refreshes, while rating credible-standard potential 4.33/5.These preferences are presented as the actionable contribution of the stakeholder consultation.
- Governance complexity: AI4Bharat’s Indic LLM Arena illustrates concentration: one organization creates benchmarks, builds models, and operates the leaderboard with government, philanthropic, and industry funding.The paper treats this concentration as characteristic of nascent ecosystems and argues that formal conflict-of-interest policies become prerequisites for trust as entities scale.
- Broader relevance: IrokoBench and AlGhafa show that the same pattern extends beyond India: regional evaluation resources exist, but trusted aggregation remains absent or concentrated.IrokoBench covers 16 African languages, while more than 2,000 African languages remain largely neglected.
7. Alternative Views
The paper argues that objections to regional leaderboards support careful, federated governance rather than abandoning governance. It emphasizes institutional accountability, asymmetric impacts, sustainable public-good infrastructure, and stronger representation.
- Governance design: Capture and gaming risks call for multi-stakeholder boards, term limits, external audits, and quarterly refreshes with contamination detection.These safeguards address local incumbent capture and leaderboard gaming without rejecting regional governance.
- Cross-regional comparison: Federation enables comparison through standardized result schemas reporting per-language and per-dialect performance side-by-side, while acknowledging that different benchmarks are not directly comparable.Identical leaderboards are unnecessary; common reporting schemas support inspection without treating scores as apples-to-apples.
- Asymmetric impact: Governance is the Global South’s only inclusion mechanism because Santhali and Hausa speakers lack the ecosystem redundancy available to Basque speakers.European-language ecosystems may have regulatory pressure, academic funding, and enterprise alternatives, while market pressure creates accountability in the Global North.
- Institutional accountability: Good intentions do not substitute for institutional accountability when expertise is concentrated and network effects create winner-takeall dynamics.The paper questions governance structures rather than organizational intentions, arguing that alternative credibility can take years to build.
- Sustainability: Evaluation infrastructure is a public good, and funding questions follow governance decisions about who bears costs, on whose behalf, and accountable to whom.NIST, MLCommons, and W3C are cited as existence proofs for publicly or collectively sustained infrastructure.
- Representation: Under transparent governance, imperfect representation can improve through user panels and complaint processes, whereas exclusion without governance is permanent.The paper accepts the representation critique as an argument for stronger governance, not against governance.
8. Call to Action
The paper calls for regional leaderboards built as publicly supported infrastructure with independent governance, transparent procedures, and sustainable funding. It also urges global maintainers, researchers, and industry to adopt stronger evaluation practices and recognize regional benchmarks as authoritative.
- Institutional design: A minimum viable regional leaderboard requires independent multi-stakeholder governance, a published COI policy, standardized submissions, dispute resolution, and interoperable result schemas.Sustainable funding could come from public funding, industry consortia with governance firewalls, or hybrid approaches.
- For nations and funding bodies: Nations and funding bodies should provide multi-year public-good funding, aggregate existing benchmarks transparently, and require independence criteria for funded evaluation efforts.The recommendation treats regional leaderboard infrastructure as a public good rather than an isolated research project.
- For global leaderboard maintainers: Global leaderboard maintainers should formalize COI disclosure and recusal, publish testing-access policies, create appeals processes, and include Global South benchmarks.They should also expand evaluation to morphologically appropriate metrics and code-switching.
- For researchers and industry: Researchers and industry should treat regional benchmarks as authoritative, require multilingual evaluation for general-capability claims, support multiple evaluation pathways, and contribute to benchmark development.This recommendation rejects treating regional benchmarks as secondary to global leaderboards.
9. Conclusion
AI leaderboards are infrastructure whose failures can leave Global South AI ecosystems operating on unreliable signals. Without governance that adopts, maintains, and evolves regional improvements, regional leaderboards remain necessary infrastructure rather than fragmentation.
- AI leaderboards shape state-of-the-art judgments and procurement and investment decisions, so regional failures propagate unreliable signals through local AI ecosystems.The paper characterizes leaderboards as infrastructure, not merely ranking tools.
- Absent conflict-of-interest policies, independent oversight, and metric-evolution mechanisms, leaderboards are structurally ill-suited to serving the Global South despite existing regional benchmarks.The stated barrier is institutional: no governance structure adopts, maintains, and evolves regional improvements.
- Reliable global evaluation would require language-family-appropriate metrics, code-switching evaluation for top-10 multilingual markets, regional-accent disaggregation, culturally grounded benchmarks, and inclusion of regional suites.Named suites include IndicSUPERB, IrokoBench, AlGhafa, and SEA-HELM.
- Until these technical and inclusive conditions are met, regional leaderboards are necessary infrastructure rather than fragmentation of the evaluation landscape.The paper frames the alternative as functional infrastructure versus continued underservice of the Global South.
Disclaimer … A.2. Survey Instrument
The paper is an argumentative position paper, not a representative empirical study, and reports an illustrative practitioner consultation whose limitations are documented. The consultation gathered expert design input through a structured survey covering governance, technical evaluation, operations, roadmap, and success criteria for an Indic AI leaderboard.
- Disclaimer: The authors state that the paper reflects their personal views and not the official positions of IIT Kharagpur, Shunya Labs, or Nasscom.The consultation is explicitly described as illustrative rather than representative.
- A.1. Purpose: The consultation gathered expert design input for an Indic AI leaderboard through Nasscom’s AI stakeholder networks between December 2025 and March 2026.The instrument was iteratively developed with industry policy experts and covered governance and architecture, technical challenges and solutions, implementation and operations, and roadmap and success.
- A.2. Survey Instrument: The survey instrument allowed multiple selections unless questions were explicitly single-select and included demographic questions with consent but no respondent-identity verification.Demographics covered stakeholder category, expertise areas, and affiliation.
- A.2. Survey Instrument: Open-ended questions invited additional suggestions, while the instrument’s language-selection question defined the starting set for the Indic AI Leaderboard.The survey also asked respondents to select among specified languages and benchmarks.
- A.2. Survey Instrument: Governance questions compared hosting models, while conflict-management questions contrasted pre-emptive exclusion with disclosure and recusal.Options included Nasscom, an academia consortium, an independent non-profit, government-driven governance, and other.
- A.2. Survey Instrument: Technical evaluation questions covered Indian ASR datasets and metrics, LLM benchmarks, semantic similarity, factuality, bias and cultural metrics, and human, LLM, or hybrid evaluation.Named ASR metrics included WER, CER, and KER; anti-gaming strategies ranged from static public to adversarial.
- A.2. Survey Instrument: Operational questions addressed submission requirements, quarterly or half-yearly cadence, emerging priorities, and success criteria including use at 2 years, procurement influence, credibility, and ASR/LLM track parity.Submission requirements included API documentation, technical specifications, reliability, and security.
A.3. Sampling and Limitations … D. Cross-Modal Governance Gap
The appendix reports statistically sharp but non-population-representative preferences from a practitioner-heavy convenience sample, supporting formal, adaptive, locally validated governance while documenting cross-modal institutional gaps.
- A.3. Sampling and Limitations: 4 of 82 respondents (4.9%) identified primarily with Government/Policy, while civil-society and end-user voices remained underrepresented.The sample was drawn from Nasscom’s networks, introducing selection bias toward practitioners engaged with the association.
- A.4. Note on Survey Evolution (n=58 to n=82): 82 final responses replaced the interim n=58 snapshot, without changing the direction of reported findings.The largest snapshot shift concerned disclosure-versus-exclusion preference, reported as 71% at n=58 before the final dataset.
- B. Selected Practitioner-Consultation Findings (n=82): The consultation’s selected findings use n = 82, with Wilson 95% confidence intervals, exact binomial or Pearson χ2 tests, and Likert t-tests against neutral mean 3.Charts generally show raw response counts, with multi-select totals explicitly exceeding 82.
- B.1. Respondent Demographics: 74% of selections came from builders, compared with 3.9% from civil society and 3.1% from government/policy, so findings describe an engaged practitioner community rather than a representative population.Production AI developers numbered 52, while Government/Policy had 4 and Civil Society/NGO had 5 structured selections.
- B.5. ASR Datasets and Metrics (Q4, Q5); B.6. LLM Datasets (Q6) and Metric Families (Q7–Q9); B.7. Evaluation Method (Q10): Vistaar (63/82, 77%) and MILU (59/82, 72%) led dataset preferences, while WER (84%), BERTScore (77%), and hybrid LLM+human evaluation (62/82, 75.6%) led metric and method choices.The results indicate demand for multi-dataset, multi-metric evaluation rather than a single canonical benchmark or score.
- B.10. Emerging Technology Priorities (Q14); B.11. Success Criteria (Q15, 1–5 Scale); B.12. What the Statistical Inference Adds: All multi-option preferences were significantly non-uniform at p < 10−4 or better, dichotomous majorities were significant at p < 10−3, and Likert means exceeded neutral at p < 10−7.These results support sharp within-sample preferences but cannot establish population-level claims because the sample was a convenience pool.
- C. Illustrative Governance Framework; D. Cross-Modal Governance Gap: The proposed governance framework is illustrative and should be adapted locally; across medicine, agriculture, and weather, known technical fixes remain blocked by missing validation mandates and appeals processes.Examples include nearly 80% misclassification by Western cardiovascular risk models in India and PlantVillage accuracy falling from 99.35% in a North American lab benchmark to 49% for cassava detection in Tanzania.
E. Federation Architecture Sketch
The proposed architecture uses federation rather than unification to enable cross-regional comparability, separating regional metric autonomy, shared reporting, and trusted-leaderboard recognition. It combines a shared result schema, a governance registry, and an optional neutral meta-view without imposing a global ranking.
- Architecture rationale: Federation, not unification, is proposed as the answer to cross-regional comparability.The appendix presents this as a starting architecture.
- Layer 1: Shared result schema: A shared JSON schema reports per-language, per-dialect, per-domain, and per-task performance while allowing regional leaderboards to weight metrics differently.The schema is descriptive rather than prescriptive about which metrics to optimize.
- Layer 2: Trusted-leaderboard registry: A trusted-leaderboard registry requires a published COI policy, documented appeals process, public methodology, and independent fairness-audit panel.Conformance is self-declared with periodic peer review.
- Layer 3: Optional meta-view: An optional meta-leaderboard hosted by a neutral body can surface cross-regional comparisons without overriding regional rankings or imposing a global ranking.The meta-view exposes capability claims to multilingual scrutiny.
- Design principles: The three layers separate what to measure, how to report measurements, and who can claim trust, while protocol specification remains future work.The architecture identifies regional autonomy, shared schema, and registry-based trust as distinct concerns.