Source-linked AI summary

Dataset Scarcity Limits Robust Evaluation of Multilingual Embedding Models: A Case Study of Slavic Languages

Ana Gjorgjevikj, Barbara Koroušić Seljak, Tome Eftimov

arXiv:2608.24477v1cs.CL

TL;DR

Multilingual embedding benchmarks provide uneven and potentially weak evidence for mid- and low-resource languages. This paper evaluates Slavic-language MTEB results with a framework combining task-specific and cross-task stability with evidence strength, finding severe benchmark sparsity alongside a small group of consistently transferable models.

  • Problem

    Multilingual embedding evaluation is uneven across languages, with limited dataset availability, task coverage, and insight into robustness under sparse conditions.

  • Method

    The paper analyzes Slavic-language MTEB benchmarks across task-specific and cross-task scopes using ranking stability, transfer consistency, and Evidence Strength Score.

  • Results

    Severe benchmark sparsity limits robustness assessment, while llama-embed-nemotron-8b, multilingual-e5-large-instruct, and Qwen3-Embedding variants consistently perform well across diverse settings.

  • Takeaways & Limitations

    Benchmark rankings and robustness conclusions should be interpreted jointly with the strength of their supporting evidence.

  • Takeaways & Limitations

    The study relies on a single snapshot of publicly reported MTEB results and does not model per-dataset uncertainty across data splits.

Abstract

from arXiv · show

Multilingual text embedding models enable cross-lingual transfer of knowledge across a wide range of NLP tasks, but their evaluation remains highly uneven across high-, mid- and low-resource languages. In this paper, we propose a two-dimensional framework, specifically tailored for analyzing multilingual embedding benchmarks under dataset scarcity, and apply it on the Slavic-language subset of the MTEB benchmark. The framework distinguishes between task-specific and cross-task evaluation, while jointly analyzing three complementary aspects: (1) ranking robustness, (2) model consistency, and (3) evidence strength. At the task-specific level, we evaluate the stability of model rankings under changes in ranking methodology and benchmark dataset composition. At the cross-task level, we assess the ability of models to generalize across diverse tasks within a language. To quantify the reliability of benchmark conclusions, we introduce an Evidence Strength Score that accounts for dataset availability, diversity, and robustness assessability. Our analysis reveals severe benchmark sparsity, with many Slavic language-task pairs relying on a single dataset or highly correlated benchmark collections, limiting the ability to draw robust conclusions. The cross-task analysis reveals a small group of highly transferable models, most notably llama-embed-nemotron-8b, multilingual-e5-large-instruct, and Qwen3-Embedding variants, that consistently perform well across Slavic languages and tasks. Overall, the results demonstrate that benchmark rankings and robustness conclusions must be interpreted jointly with certain notation of their evidence strength and highlight benchmark scarcity as a major obstacle to trustworthy multilingual evaluation.

1 Introduction

The paper introduces a two-dimensional framework for evaluating multilingual embedding benchmarks under dataset scarcity and applies it to Slavic languages in MTEB. It combines ranking robustness, model consistency, and evidence strength to show that sparse evidence limits trustworthy conclusions despite a small group of consistently transferable models.

  • Framework: The framework evaluates multilingual benchmarks at task-specific and cross-task levels under dataset scarcity and imbalance.Task-specific analysis examines ranking stability and top-k transfer consistency within a language and task, while cross-task analysis assesses generalization across task families.
  • Framework: It jointly analyzes ranking stability, top-k transfer consistency, and evidence strength across aggregation methods, dataset compositions, and evaluation conditions.Evidence strength quantifies how much confidence can be placed in benchmark conclusions, while the other components assess ranking stability and model consistency.
  • Contribution: Evidence Strength Score is the paper’s novel contribution, addressing low-resource multilingual evaluation through qualitative evidence levels and a quantitative reliability measure.The score characterizes which stability analyses are feasible for each task–language pair and how reliable the available benchmark evidence is.
  • Findings: Slavic-language MTEB evaluation reveals severe benchmark sparsity and redundancy, with most language–task pairs relying on a single dataset or lacking diversity for reliable robustness assessment.These limitations make apparently stable benchmark conclusions difficult to interpret confidently.
  • Findings: A small group of large multilingual models—llama-embed-nemotron-8b, multilingual-e5-large-instruct, and Qwen3-Embedding variants—consistently demonstrates stable cross-task performance.The paper identifies these models as the most notable examples despite the broader benchmark limitations.

2 Related Work

Prior work established multilingual embedding models and broad benchmarks, but related research identifies persistent limitations in language-specific coverage, evaluation stability, and evidence-aware analysis. These gaps motivate jointly examining ranking stability and cross-task consistency under uneven and sparse data.

  • Multilingual text embedding models: Multilingual embedding models map multiple languages into a shared space for cross-lingual retrieval and semantic similarity, with newer systems using contrastive training and foundation-model architectures.Examples include LaBSE, multilingual Sentence-BERT, multilingual E5, Qwen3, and LLaMA-based embedding models.
  • Multilingual embedding benchmarks: MTEB unified multilingual evaluation, while MMTEB expanded coverage to over 500 datasets and 250 languages; aggregate scores still limit language-specific analysis.Language-focused benchmarks such as PL-MTEB and SEB expose performance differences that broad aggregates can obscure.
  • Evaluation stability: Benchmark results can vary with input changes, heterogeneous metrics, dataset imbalance, and correlations, while alternative ranking methods retain assumptions about equal dataset importance.These limitations challenge simple averaging and complicate interpretation of leaderboard differences.
  • Low-resource languages and evaluation gaps: Low-resource languages, including Slavic languages, remain underrepresented, with existing resources often limited to particular languages or tasks.Languages spoken by smaller populations are especially underrepresented.
  • Summary: Prior work leaves unaddressed whether sparse or redundant evidence supports stability conclusions and lacks a joint analysis of task-specific ranking stability and cross-task transfer consistency.The identified gaps concern unequal evaluability across task-language pairs and the absence of explicit evidence qualification.

3 Methodology

The methodology introduces a two-dimensional framework for multilingual embedding benchmarks under dataset scarcity, separating task-specific ranking analysis from cross-task top-k transfer analysis. It jointly evaluates ranking robustness, model consistency, and evidence strength to calibrate confidence in benchmark conclusions.

  • Two-dimensional framework: The framework separates task-specific ranking stability within a language-task pair from cross-task top-k transfer consistency across tasks within a language.Task-specific analysis examines ranking stability under aggregation methods and dataset compositions, whereas cross-task analysis evaluates model transfer consistency.
  • Two-dimensional framework: Ranking stability is excluded from cross-task analysis because scores from retrieval, clustering, and classification reflect different objectives and metrics.Cross-task score disagreement may therefore indicate task specialization rather than instability.
  • Model consistency: The top-k transfer measure extends prior binary membership analysis by weighting models according to their actual position within the top-k set.This complements ranking stability, which compares complete rankings across ranking schemes and decorrelated dataset compositions.
  • Evidence strength: Evidence strength is introduced as a novel contribution through qualitative evidence levels and the quantitative Evidence Strength Score (ESS).The ESS combines dataset availability, effective diversity, aggregation-stability assessability, and composition-stability assessability, with equal weighting.
  • Evidence strength: Five or more datasets are treated as saturated availability in ESS, with nmax = 5 because additional datasets yield diminishing returns in evidence strength.Effective diversity separately captures redundancy, with values near 1 indicating largely independent evaluation evidence.

4 Experimental Design

The study evaluates open, fully zero-shot embedding models on eight MTEB task categories using multiple dataset compositions and diverse ranking methods to assess evaluation robustness.

  • Experimental data: The analysis uses MTEB Multilingual leaderboard v21 data across eight task categories, restricted to open, fully zero-shot models.The categories include classification, clustering, retrieval, reranking, semantic textual similarity, pair classification, multilabel classification, and bitext mining.
  • Dataset compositions: Datasets with pairwise correlation above τ = 0.9 are clustered, and decorrelated subsets sample one dataset per cluster.The procedure is repeated three times, producing one all-dataset composition and three decorrelated compositions.
  • Ranking methods: Model rankings are computed for each language, task, and dataset composition using WSM, TOPSIS, VIKOR, and PROMETHEE II.PROMETHEE II uses both usual and Gaussian preference functions, covering value-based, distance-based, and compromise-based ranking logic.

5 Results

The results show that Slavic multilingual embedding evaluation is highly uneven and sparse, especially across languages and tasks. Where stability is assessable, rankings are generally concordant, but the strongest conclusions come from settings with stronger evidence.

  • Dataset coverage: Russian covers 100% of tasks, while Ukrainian covers 37.5% and Bosnian and Belarusian each cover 25.0%.Bitext mining contains 3–5 datasets per language for most languages, whereas several other tasks are underrepresented or absent.
  • Evidence quality: STS, reranking, and pair classification are dominated by absent or single-dataset evidence, while correlated datasets can make multiple resources effectively redundant.Benchmark sparsity and redundancy limit reliable stability analysis, particularly for lower-coverage Slavic languages.
  • Evidence strength: 0.85 is Russian’s highest ESS for bitext mining, and Russian and Polish form the strongest evidence cluster with broader coverage and more diverse resources.Czech, Serbian, and Croatian form an intermediate cluster, while Bosnian, Slovak, Slovene, and Macedonian form the weakest group.
  • Ranking stability: 0.990 is the mean W_RS across assessable cases, indicating that aggregation methods generally produce highly similar model rankings.For robust bitext-mining compositions, mean W_rob RS = 0.970 and mean W_rob DS = 0.983, showing limited sensitivity to dataset perturbations.
  • Evidence and stability: 0.829 is the average ESS for ERS+DS cases, compared with 0.615 for ERS, 0.300 for E1, and 0.232 for ESC.Stability is therefore most assessable where benchmark evidence is strongest, while benchmark scarcity—not ranking instability—remains the primary limitation.
  • Classification transfer consistency: 0.70 is Polish classification’s ESS, supported by four independent datasets (n = q = 4) and enabling aggregation stability analysis (ERS).Classification is otherwise sparse, with most languages supported by E1 or E0 and ESS ≤0.30.

6 Discussion

The discussion distinguishes task-specific superiority from cross-task generalization: task-specialist leaders often do not transfer broadly, while a small set of models repeatedly performs strongly across tasks. Cross-task rankings are more stable than task-specific rankings, indicating task specialization is a major source of evaluation variability.

  • Cross-task generalization: Task-specific winners frequently differ from the strongest cross-task models, showing that narrow-domain superiority does not ensure broad generalization.The discussion contrasts exceptional performance within individual domains with weaker representation among the strongest cross-task models.
  • Task specialization: Classification and pair classification favor Qwen3-Embedding variants, while clustering and reranking favor llama-embed-nemotron-8b across most evaluated languages.Retrieval, STS, multilabel classification, and bitext mining have different leaders, further illustrating task specialization.
  • Task specialization: No single architecture dominates all embedding use cases: LaBSE-ru-turbo, Octen-Embedding-8B, KaLM-Embedding-Gemma3-12B-2511, and bge-m3 lead specific task families.The passage also identifies multilingual-e5-large-instruct as a bitext-mining leader under some evidence regimes.
  • Cross-task generalization: Llama-embed-nemotron-8b’s cross-task dominance is partly supported by broad clustering-benchmark availability and consistent strength on that well-covered task family.Because cross-task consistency aggregates evaluable tasks, strong performance on a widely available task family contributes substantially to final rankings.
  • Ranking stability: Cross-task rankings repeatedly highlight the same small model set, whereas task-specific leaders vary by language, dataset composition, and task.This greater cross-task stability suggests task specialization, rather than language variation, is the principal source of evaluation variability.
  • Specialists and generalists: The findings separate task specialists from generalists: specialists excel in specific task families, while llama-embed-nemotron-8b, multilingual-e5-large-instruct, and Qwen3-Embedding variants are identified as generalists.The specialist examples include LaBSE-ru-turbo, Octen-Embedding-8B, KaLM-Embedding-Gemma3-12B-2511, and bge-m3.

7 Conclusion

The study finds that dataset scarcity constrains reliable multilingual embedding evaluation across Slavic languages. Despite these limitations, a small group of large multilingual models shows stable cross-task transfer consistency, while current benchmarks remain insufficient for many languages.

  • Conclusion: Dataset scarcity largely constrains evaluation across Slavic languages and limits the reliability of benchmark conclusions.The analysis considers both task-specific stability and cross-task transfer consistency.
  • Conclusion: A small group of large multilingual models, including llama-embed-nemotron-8b, Qwen3-Embedding, and multilingual-e5, consistently demonstrates stable cross-task transfer consistency.Most other models remain task- and language-specific.
  • Conclusion: Current benchmarks provide insufficient evidence for reliable conclusions in many Slavic languages, underscoring the need for richer and more task-balanced evaluation resources.The conclusion links benchmark insufficiency to the scarcity of available datasets.

Limitations · A Evidence Strength Design Choices

The paper frames ESS as a confidence measure for benchmark evidence, not model performance, and designs it to reflect dataset availability, diversity, and stability assessability. Its limitations include reliance on a single snapshot of publicly reported MTEB results and the risk that sparse or redundant benchmarks produce apparently stable rankings.

  • Limitations: The study analyzes how benchmark conclusions change with dataset composition and aggregation methods rather than selecting models.It complements aggregate metrics by assessing benchmark-result stability.
  • Limitations: The analysis relies on a single snapshot of publicly reported MTEB results, which may change over time.
  • A Evidence Strength Design Choices: ESS quantifies the reliability of available benchmark evidence rather than model performance itself.Its motivation is that sparsity, dataset redundancy, and limited stability assessability affect stability conclusions in low-resource multilingual settings.
  • A Evidence Strength Design Choices: Dataset availability is normalized with a saturation threshold nmax = 5 to model diminishing returns and prevent high-resource languages from dominating through dataset quantity.
  • A Evidence Strength Design Choices: Effective dataset diversity is the ratio of decorrelated dataset clusters to total datasets, distinguishing independent evidence from highly correlated evaluations.Datasets are clustered using similarities in model-performance profiles.
  • A Evidence Strength Design Choices: Aggregation stability and composition stability measure whether multiple ranking schemes and multiple non-redundant dataset compositions can be compared.Together, they quantify whether stability analysis is feasible for a task–language pair.
  • A Evidence Strength Design Choices: ESS combines four components with equal weighting to preserve interpretability and avoid assumptions about the relative importance of quantity, diversity, and assessability.The resulting score is a normalized estimate of benchmark-conclusion reliability for a task–language pair.
  • A Evidence Strength Design Choices: ESS complements rather than replaces stability metrics, indicating whether stable conclusions are supported by sufficient and diverse evidence.High stability with low ESS may reflect sparse or redundant evaluation conditions.

B Interpretation of ESS Values

ESS is a confidence indicator for the strength of stability evidence, not a direct measure of model quality. Its interpretation accounts for benchmark coverage, dataset diversity, and stability assessability, while penalizing missing or redundant evaluation evidence.

  • B Interpretation of ESS Values: ESS values below 0.1 indicate sparse or missing benchmark coverage, warranting substantial caution in interpreting stability conclusions.Scores from 0.1 to 0.25 indicate weak evidence, typically caused by isolated datasets or limited stability assessability.
  • B Interpretation of ESS Values: ESS scores between 0.25 and 0.4 represent moderate evidence because partial stability analysis is feasible despite incomplete benchmark coverage.Higher ESS values indicate that stability and transfer-consistency results are supported by broader, more reliable evaluation evidence.
  • B Interpretation of ESS Values: Russian achieves the highest language-level ESS at 0.49, representing comparatively strong evidence within the current benchmark landscape rather than near-saturated evaluation.The ESS formulation penalizes missing tasks, dataset redundancy, and limited stability assessability, so even strong-performing languages may receive only moderate values.

C Language-Specific Cross-Task Analysis

The language-specific cross-task analysis evaluates transfer consistency across all eight tasks for each language, primarily measuring stability across aggregation methods. Using k = 10, it compares ranking schemes and coverage weighted aggregation, with results visualized as model-by-ranking-scheme dendrograms.

  • Method: Transfer consistency is evaluated across all eight tasks for each language.The analysis is language-specific and focuses on cross-task transfer.
  • Method: Because not all tasks support composition stability analysis, conclusions primarily concern stability across aggregation methods.The reported scope is therefore narrower than full composition-stability evaluation.
  • Method: With k = 10, cross-task transfer consistency is computed by ranking scheme and coverage weighted aggregation.The resulting consistencies are visualized as dendrograms whose rows represent models and columns represent ranking schemes.

C.1 Western South Slavic Languages · C.2 Eastern South Slavic Languages

Western and Eastern South Slavic results identify highly consistent multilingual embedding models, but the strength of conclusions varies with benchmark coverage. Slovenian evidence is limited, whereas Bulgarian supports a stronger cross-task generalization analysis with a clear leading pair.

  • C.1 Western South Slavic Languages: Slovenian has benchmark data for only five of the eight analyzed tasks.
  • C.1 Western South Slavic Languages: multilingual-e5-large-instruct achieves the highest Slovenian coverage-weighted consistency (CW = 0.35), followed closely by llama-embed-nemotron-8b (CW = 0.34).
  • C.1 Western South Slavic Languages: A small group of large multilingual models consistently outperforms the remaining Slovenian candidates regardless of ranking methodology.
  • C.1 Western South Slavic Languages: Slovenian language-level evidence remains limited (ESS = 0.18), requiring cautious interpretation.
  • C.2 Eastern South Slavic Languages: Bulgarian has benchmark data for seven of the eight analyzed tasks, providing one of the strongest settings for cross-task generalization analysis.
  • C.2 Eastern South Slavic Languages: llama-embed-nemotron-8b leads Bulgarian consistency (CW = 0.55), narrowly ahead of multilingual-e5-large-instruct (CW = 0.53).The two models form a clear leading pair and achieve nearly identical coverage-weighted consistency scores, substantially ahead of the remaining models.
  • C.2 Eastern South Slavic Languages: Qwen3-Embedding-8B forms a strong Bulgarian second tier (CW = 0.42), followed by Octen-.

C.3 West Slavic Languages

Polish offers one of the strongest settings for evaluating cross-task transfer consistency, with benchmark data available for seven of eight analyzed tasks. Qwen3-Embedding-4B leads coverage-weighted consistency, followed by llamaembed-nemotron-8b and Qwen3-Embedding-8B.

  • Polish: Polish has benchmark data for seven of the eight analyzed tasks, making it one of the strongest settings for evaluating cross-task transfer consistency.Its broad task coverage supports cross-task analysis, although the passage does not specify the eighth task.
  • Polish: Qwen3-Embedding-4B achieves the highest coverage-weighted consistency (CW = 0.54), followed by llamaembed-nemotron-8b and Qwen3-Embedding-8B (both CW = 0.50).Octen-Embedding-8B completes the leading group with CW = 0.44.
  • Polish: multilingual-e5-large-instruct forms a second tier (CW = 0.33), while the remaining models show substantially lower cross-task consistency.The passage identifies a clear separation between the leading group, the second tier, and the remaining models.

C.4 East Slavic Languages

Russian has complete benchmark coverage across all eight analyzed tasks, providing the strongest evidence for cross-task transfer consistency. Llama-embed-nemotron-8b leads, and the conclusions have very strong evidence strength with an ESS score of 0.49.

  • Russian: Russian is the only language with complete benchmark coverage across all eight analyzed tasks, enabling the strongest cross-task transfer analysis.This coverage provides the strongest evidence for evaluating cross-task transfer consistency.
  • Russian: Llama-embed-nemotron-8b leads Russian cross-task transfer consistency with CW = 0.70, followed by Qwen3-Embedding-8B (CW = 0.65) and Octen-Embedding-8B (CW = 0.61).Qwen3-Embedding-4B (CW = 0.48) completes the leading group of highly transfer-consistent models.
  • Russian: Qwen3-Embedding-4B remains a stable fourth-ranked model, with a clear separation between leading multilingual embedding models and remaining approaches.The results indicate strong and reliable cross-task transfer consistency for Russian.
  • Russian: The ESS score of 0.49 corresponds to a very strong level of evidence and confidence in the robustness of the Russian conclusions.The evidence strength is supported by complete benchmark coverage and the observed ranking separation.
Loading 2608.24477v1…