Source-linked AI summary

From Global Benchmarks to Local Evaluations: Benchmarking LLMs for the German Public Sector

Camilla Dalerci, Thilo Michael, Robin Schaefer, Daniel Weinland

arXiv:2608.17827v1cs.CL

TL;DR

Public institutions must choose among heterogeneous LLMs whose deployment suitability extends beyond generic task performance. MÖVE evaluates German public-sector models across governance and contextual dimensions, finding trade-offs with no single model strongest across all criteria. This supports context-specific evaluation rather than universal performance rankings.

  • Problem

    Public institutions choose among heterogeneous models differing in performance, energy requirements, documentation practices, and contextual knowledge, while suitability extends beyond generic task performance.

  • Method

    MÖVE evaluates 39 open-weight and proprietary models across task performance, provider transparency, energy consumption, and knowledge of German political-party positions.

  • Results

    No single model is strongest across all dimensions; estimated energy consumption varies 63-fold, transparency is weakest for several governance disclosures, and no regional provider group consistently dominates political knowledge.

  • Takeaways & Limitations

    Public administrations should compare models against deployment-specific governance and domain requirements rather than rely on model size, provider reputation, or general benchmark performance alone.

  • Takeaways & Limitations

    The study covers only three evaluation dimensions and a fixed set of models.

Abstract

from arXiv · show

Public institutions face a persistent challenge in selecting LLMs suited to their specific context. Existing benchmarks, however, are of limited use as they primarily reflect English-language and US-centric settings, and often only evaluate task performance. In this paper, we present first results of MÖVE, a holistic evaluation framework for the German public sector, examining three rarely considered governance dimensions: energy consumption, provider transparency, and knowledge of German-party positions. Our results reveal significant trade-offs, with no single model excelling across all dimensions: estimated energy consumption varies more than 60-fold and is not explained by model size alone, information disclosure varies systematically across providers, and European models do not exhibit stronger knowledge of German party positions. Model selection for public institutions thus cannot rely on performance rankings alone. Instead, evaluations should also reflect the governance requirements of the deployment context.

1 Introduction

Public institutions need locally grounded LLM evaluations because general benchmarks offer limited evidence for suitability in specific institutional, linguistic, or societal contexts. MÖVE addresses this gap by evaluating 39 LLMs across three governance dimensions for the German public sector.

  • Motivation: Public institutions now choose among heterogeneous LLMs differing in size, performance, energy requirements, documentation practices, and contextual knowledge.This expanding model landscape makes selection more complex than choosing between one or two established providers.
  • Motivation: General benchmark performance alone offers limited evidence about suitability for a particular institutional, linguistic, or societal context.Identifying potential harms and anticipating failures in intended use cases make model evaluation essential.
  • Benchmark limitations: Established benchmarks predominantly measure task performance on English-language datasets and often reflect US- or UK-specific contexts.Translated multilingual datasets frequently fail to address the underlying domain or context gap.
  • MÖVE framework: MÖVE evaluates 39 LLMs across energy consumption, provider transparency, and political knowledge for the German public sector.The framework estimates inference-time energy from task-specific output statistics, assesses transparency with 21 questions across seven documentation domains, and evaluates political knowledge using 4,788 official positions from 64 German political parties.

2 Benchmark Design

MÖVE evaluates 39 LLMs from 13 providers across three governance dimensions relevant to German public-sector deployment. Its stakeholder-informed design combines energy, transparency, and politically grounded factual-knowledge assessments.

  • Benchmark scope: 39 LLMs from 13 providers are evaluated across inference-time energy consumption, provider transparency, and knowledge of German political-party positions.The benchmark was developed through semistructured feedback sessions with public-sector representatives and exchanges with practitioners.
  • Energy consumption: Energy consumption is estimated with EcoLogits using model parameters and actual output-token counts from nine German-language datasets.Results are reported as mean estimated energy consumption in watt-hours per inference request, by task and across tasks; proprietary models without disclosed parameter counts use EcoLogits assumptions.
  • Provider transparency: Provider transparency is measured with 21 questions across seven domains, using publicly verifiable information from provider websites, model cards, and technical publications.Questions are scored 0 for no information, 1 for partial information, and 2 for clear and verifiable information; EU General-Purpose AI Code of Practice endorsement is binary.
  • Provider transparency: 27.4% of transparency-assessment items initially diverged, while explicit scoring notes reduced divergence to 16.2% across a protocol whose complete output was reviewed by a second researcher.The 16.2% figure measures final protocol internal consistency rather than accuracy; across 39 models, the dataset contains 819 model-question assessments.
  • Political knowledge: Political knowledge is evaluated against 4,788 official Wahl-O-Mat positions from 64 parties across the 2013, 2017, 2021, and 2025 federal elections.Models predict each specified party’s position on German-language propositions, and performance is measured by classification accuracy against the official responses.

3 Results

The evaluated models exhibit substantial trade-offs across energy consumption, provider transparency, and knowledge of German political-party positions. No model consistently performs best across these governance dimensions.

  • Energy Consumption: 63-fold variation in estimated mean energy consumption spans 0.647 Wh to 40.6 Wh per query across 39 models, with a median of 1.788 Wh.Smaller models generally consume less energy, but model size does not fully determine efficiency; proprietary-model estimates require caution because parameter counts are undisclosed.
  • Energy Consumption: Task and model characteristics affect energy use: summarization is more energy-intensive than question answering and topic extraction, while reasoning models incur particularly high costs.Longer outputs help explain summarization’s higher energy requirements.
  • Provider Transparency: 11.5% is the mean score for compute and energy transparency; 84.6% disclose no training-energy consumption and 79.5% provide no measurement methodology.Providers signing the EU General-Purpose AI Code of Practice score only marginally higher overall than non-signatories.
  • Political Knowledge: 0.671 is the highest accuracy achieved for German party-position classification across 4,788 positions from 64 parties, and no model is consistently accurate.Larger models tend to perform better, but size does not determine performance: GPT-4o Mini matches DeepSeek R1 at the top, while medium-sized and fine-tuned models remain competitive.

4 Discussion

The discussion argues that public-sector model selection must move beyond performance rankings toward multidimensional, locally grounded evaluations. Energy costs, provider transparency, and local political knowledge each reveal limitations of common suitability proxies.

  • Energy consumption: 63-fold variation in estimated energy consumption shows that resource efficiency is essential because similarly performing models can impose substantially different operational and environmental costs.Energy should receive particular attention in high-volume public-sector applications.
  • Provider transparency: Providers disclose less about training data, bias mitigation, computational resources, and energy consumption than about basic model properties, limiting independent risk assessment and evidence-based procurement.The findings identify a persistent accountability gap in provider transparency.
  • Political knowledge: European, US, and Chinese models achieve competitive results on German political-party positions, so provider geography is not a reliable proxy for local suitability.No regional group consistently dominates, while the modest maximum accuracy indicates that all models require evalu…
  • Implications: Model size, provider reputation, and general benchmark performance alone are insufficient; public administrations should compare models against specific governance and domain requirements.The conclusion calls for multidimensional, locally grounded evaluation and stronger transparency requirements from providers and regulators, replacing universal rankings with deployment-context evaluations.

Limitations

The study evaluates only three dimensions and a fixed model set, using comparative energy estimates and documentation-based transparency scores. Its datasets cover selected aspects of the German public-sector context rather than the full diversity of tasks, political knowledge, or deployment conditions.

  • Scope: The study is limited to three evaluation dimensions and a fixed set of models, with future work planned to add newer models and criteria.This expansion aims to provide a more comprehensive assessment of public-sector suitability.
  • Measurement: Energy results are comparative estimates, while proprietary-model estimates rely on EcoLogits’ assumptions about undisclosed parameter counts.Transparency scores reflect publicly available documentation collected and verified between March and May 2026.
  • Representativeness: The datasets capture selected aspects of the German public-sector context and cannot represent the full diversity of administrative tasks, political knowledge, or deployment conditions.

A Transparency Scores

Transparency scores across all 39 models reveal a consistent divide between well-documented and poorly documented domains. The strongest documentation concerns model identification, architecture, and distribution & access, while use & deployment, training & data, and compute & energy lag behind.

  • A Transparency Scores: Transparency scores cover all 39 models and are sorted by domain and total score.Each model’s score is presented across the seven transparency documentation domains.
  • A Transparency Scores: Seven transparency domains are represented, with each segment corresponding to one domain’s score.
  • A Transparency Scores: A pronounced gap separates well-documented domains—Model Identification, Architecture, and Distribution & Access—from poorly documented ones.The poorly documented domains are Use & Deployment, Training & Data, and Compute & Energy, and this pattern appears across nearly all models.
Loading 2608.17827v1…