Source-linked AI summary
compar:IA: The French Government's LLM arena to collect French-language human prompts and preference data
Lucie Termignon, Simonas Zilinskas, Hadrien Pélissier, Aurélien Barrot, Nicolas Chesnais, Elie Gavoty
TL;DR
LLMs and their alignment data remain disproportionately English, leaving open preference resources scarce for other languages. compar:IA addresses this gap with an accessible French public LLM arena using blind pairwise comparisons and open data release, collecting more than 600,000 prompts and 250,000 votes by 2026-02-07. The platform provides reusable infrastructure for multilingual evaluation and human-AI interaction research, while its arena setting and pairwise design limit representativeness and absolute-quality assessment.
Problem
French and other non-English languages receive limited training and open human preference data, despite preference data’s role in RLHF and DPO.
Method
compar:IA uses an accessible public arena with blind pairwise model comparisons, unconstrained prompts, open data publication, and privacy filtering.
Results
600,000+ free-form prompts and 250,000+ preference votes were collected by 2026-02-07, while analyses characterized French user interactions across major use types.
Takeaways & Limitations
compar:IA demonstrates that accessible public-sector interfaces can collect large-scale human evaluation and open non-English preference data from general users.
Takeaways & Limitations
Arena users may submit shorter, test-oriented prompts, and pairwise comparisons may miss absolute quality, long-term usefulness, factual completeness, or safety dimensions.
Abstract
from arXiv · showhide
Large Language Models (LLMs) often show reduced performance, cultural alignment, and safety robustness in non-English languages, partly because English dominates both pre-training data and human preference alignment datasets. Training methods like Reinforcement Learning from Human Feedback (RLHF) and Direct Preference Optimization (DPO) require human preference data, which remains scarce and largely non-public for many languages beyond English. To address this gap, we introduce compar:IA, an open-source digital public service developed inside the French government and designed to collect large-scale human preference data from a predominantly French-speaking general audience. The platform uses a blind pairwise comparison interface to capture unconstrained, real-world prompts and user judgments across a diverse set of language models, while maintaining low participation friction and privacy-preserving automated filtering. As of 2026-02-07, compar:IA has collected over 600,000 free-form prompts and 250,000 preference votes, with approximately 89% of the data in French. We release three complementary datasets -- conversations, votes, and reactions -- under open licenses, and present initial analyses, including a French-language model leaderboard and user interaction patterns. Beyond the French context, compar:IA is evolving toward an international digital public good, offering reusable infrastructure for multilingual model training, evaluation, and the study of human-AI interaction.
1 Introduction
compar:IA addresses the scarcity of open, non-English preference data by collecting French-language prompts and judgments through an accessible public LLM arena. By 2026-02-07, it had accumulated substantial openly released data and begun supporting broader multilingual use.
- French typically represents a very small share of LLM pre-training and post-training data, limiting language coverage in model development.Llama 2 reports French at 0.16% of its training corpus.
- Non-English LLMs can exhibit degraded fluency, mismatched register, culturally inappropriate responses, and weaker safety guarantees.
- Large-scale open preference datasets remain rare outside English, despite their importance for RLHF and DPO.
- compar:IA adapts blind pairwise comparison for a predominantly French-speaking general audience by lowering participation barriers and explaining the process.
- 600,000+ free-form prompts and 250,000+ preference votes had been collected by 2026-02-07, with continuously released data and privacy filtering.The platform also began adapting its interface to other languages in November 2025.
- The paper documents compar:IA’s platform design, datasets, adoption indicators, limitations, and future multilingual expansion.
2 The compar:IA platform
The compar:IA platform combines unconstrained prompts, blind side-by-side model responses, layered feedback, and low-friction access. Its public-facing design emphasizes accessibility, model diversity, user awareness, and progressively more stable infrastructure.
- User interaction flow: Users enter unconstrained free-form prompts, receive two anonymous side-by-side responses, and may continue into multi-turn interactions.
- User interaction flow: Users provide complementary feedback through message-level reactions or conversation-level preference votes.
- User interaction flow: After feedback, the platform reveals model identities, provides metadata, and estimates inference energy use from generated tokens and model characteristics.
- Accessibility and participation: No account is required, and contextual explanations and visual cues support participation by non-expert users.The low-friction baseline limits metadata collection but is intended to broaden access.
- Value for end users: 104 models were available as of 2026-02-07, including 29 proprietary models and many open-weight or open-source systems.
- Value for end users: Side-by-side comparisons expose linguistic, stylistic, and cultural differences between models, while energy estimates introduce an environmental dimension of use.
- Technical evolution: The platform evolved from a Gradio-based prototype into a FastAPI and SvelteKit service with self-financed, per-token inference.
3 Data collection and dataset construction
compar:IA has continuously collected predominantly French conversational data through open, privacy-filtered datasets, including prompts, preferences, reactions, and energy estimates.
- Collected data: Over 600,000 free-form prompts and more than 250,000 conversation-level preference votes and message-level reactions had been collected by 2026-02-07.The platform opened publicly in October 2024 and has collected conversation data continuously since then.
- Language distribution: 89.14% of collected prompts are in French, although the platform is technically multilingual.French accounts for the large majority of the data.
- Prompt categories: Technical and educational prompts constitute 31.95% of the distribution through the combined Natural Science and Education categories.Other topics are represented as well, though educational-sector usage may be overrepresented.
- Dataset structure: The data are published as conversations, conversation-level votes, and message-level reactions, hosted on Hugging Face and mirrored on data.gouv.fr.The datasets are released under the Etalab 2.0 open license.
- Privacy filtering: Before publication, an LLM-based detector filters conversations containing personal or sensitive information, excluding about 5% of conversations with associated votes and reactions.The conservative policy excludes entire conversations rather than masking individual text spans, prioritizing privacy over dataset size.
- Comparison and enrichment: Compared with existing open preference data, compar:IA provides several hundred thousand French-language prompts and interactions, while French comprised 1.5% of LMSYS Chat-1M.The datasets also include electricity-consumption estimates calculated with the Ecologits method.
4 Adoption, usage, and impact
compar:IA has attracted sustained voluntary public participation and supports educational, analytical, and evaluative outputs, including a preference-based model leaderboard. These indicators suggest active interest, but reuse and leaderboard interpretation remain limited by representativeness and self-selection.
- Adoption: More than 300,000 unique visitors used compar:IA by 2026-02-07, with continuous organic traffic rather than campaign-specific participation.Participation is voluntary, unpaid, and does not require accounts.
- Dataset reuse: 778 unique users requested access to the three datasets on Hugging Face, while the data.gouv.fr copies are not gated.A survey of 25 respondents reported uses primarily for model training, research, and evaluation.
- Limitations: Systematic reuse reporting remains limited, particularly among industrial actors, making impact measurement challenging.Leaderboard positions are also influenced by prompt distribution, user population, and self-selection effects.
- Education: The platform has been integrated into workshops, talks, and digital-literacy programs, with more than 1,400 potential facilitators registering for Les Duels de l’IA materials.PIX expects more than 1.5 million students to use compar:IA through its 2026 AI curriculum.
- Model leaderboard: compar:IA released a model leaderboard based on aggregated pairwise preferences, with rankings updated weekly.The leaderboard was developed with PEReN and made public in November 2025.
- Leaderboard methodology: The leaderboard uses Bradley–Terry models on conversation-level votes and message-level reactions to represent relative user preferences rather than task-specific performance.Its primary role is exploratory and educational, not a formal benchmark.
- Thematic analysis: Analysis of over 175,000 conversations identified learning, advice seeking, content generation, and information retrieval as four dominant interaction types.Health prompts were predominantly advice-oriented, scientific topics learning-focused, and creative domains emphasized content generation.
5 Use cases for the AI ecosystem
The released compar:IA datasets support preference-based training, synthetic-data generation, usage research, and multilingual evaluation. Their central value is providing real-world French prompts and preferences for research and model-development workflows.
- Training and development: The datasets are primarily intended for research and model-development workflows using human prompts and preference data.They are directly applicable to reinforcement learning from human feedback and direct preference optimization.
- Synthetic data and usage research: Prompts and conversations can seed controlled synthetic-data generation for additional training data.Prompt collections also support analysis of real-world usage distributions, topic prevalence, and interaction styles.
- Evaluation: Sampling prompts and preferences enables multilingual evaluation benchmarks grounded in actual user behavior, especially for underrepresented languages.This provides an alternative to relying solely on expert-designed tasks.
6 Governance and development history
compar:IA developed as a non-commercial French public service that combines public awareness of LLMs with open collection of French human prompt and preference data. It evolved iteratively through public-sector collaboration and is expanding toward a multilingual digital common.
- compar:IA originated within French public administration to improve access to French-language data for language-model training while respecting copyright.
- The platform was developed through workshops, prototype testing, and iterative improvements before becoming a stable national digital public service.
- compar:IA is jointly operated by the Ministry of Culture and DINUM as a non-commercial digital public service.
- Its objectives are raising awareness of LLM diversity, biases, and environmental impacts while collecting French prompts and preferences for open release.
- Since November 2025, compar:IA has been recognised as a digital public good, with free open-source infrastructure and openly licensed datasets.
- Expansion to other European languages currently uses bilateral partnerships, with a medium-term aim of shared governance for a digital common.
7 Limitations
The platform’s open, voluntary arena design creates important limits on representativeness, task coverage, preference interpretation, and model comparability. These constraints should temper how its data and leaderboard are used.
- compar:IA cannot characterize users demographically or apply weighting schemes to correct population imbalances because it does not collect socio-demographic information.
- Professional, confidential, and regulated use cases are likely underrepresented, limiting prompt diversity and applicability to domains such as law and healthcare.
- Leaderboard preferences from general-public usage may not reflect professional performance and should not be treated as task-specific evaluations.
- Voluntary self-selection and outreach through educational networks may bias preferences toward people already interested in AI or digital tools.
- Arena users may adopt an evaluative mindset and submit shorter, simplified prompts, while French-oriented framing may overrepresent culture-specific questions.
- Pairwise comparison captures relative differences but can obscure absolute quality, factual completeness, long-term usefulness, or safety considerations.
- Historical system-prompt asymmetries, opaque proprietary pipelines, and undisclosed quantization can confound model comparisons and leaderboard rankings.
- Preference signals can reflect stylistic, latency, and voting-behavior biases, which introduce noise and complicate interpretation despite open release enabling post hoc study.
8 Future directions
Future work extends compar:IA across languages, adds privacy-preserving context, and develops specialized arenas for professional use. These directions aim to support cross-linguistic analysis and more task-specific preference data.
- The language-agnostic architecture is being extended beyond French, initially toward underrepresented European languages and potentially other scarce-data contexts.
- Multilingual deployment could enable consistent cross-linguistic studies of model behavior and user preferences across cultural contexts.
- Optional, privacy-preserving metadata could provide contextualized preference signals and basic stratified analyses without fine-grained profiling.
- Specialized professional arenas could use controlled cohorts, explicit consent, and stricter access conditions.
- Such arenas would support higher-quality, task-specific preference data and more reliable rankings within defined domains.
9 Conclusion
compar:IA demonstrates a public-sector approach to collecting large-scale French human preference data through accessible blind pairwise evaluation and open licensing. The platform offers a replicable, multilingual-oriented model for human-centered AI evaluation.
- More than 600,000 prompts and 250,000 preference votes were released under open licenses through compar:IA’s French-language public arena.
- Accessible interfaces and minimized participation friction enabled large-scale evaluation with a general public using real-world prompts.
- Operating outside commercial incentives, the platform supports collective learning while balancing participation, transparency, and privacy.
- compar:IA provides a replicable model for language-specific, human-centered evaluation that complements English-centric benchmarks.