Source-linked AI summary
MameLoshnLM: Yiddish Language Model and Evaluation Benchmark
Uri Katz, Omer Goldman, Tomasz Limisiewicz, Reut Tsarfaty, Noah A. Smith
TL;DR
Yiddish language modeling is hindered by limited digital presence and unreliable evaluation resources. The paper introduces a curated corpus, benchmark, and Yiddish-specific model, which achieves the best overall performance across 14 evaluations and better matches native lexical and morphological patterns.
Problem
Yiddish has limited contemporary digital text and common multilingual resources contain sparse, noisy, or poorly classified material, while reliable evaluation resources remain scarce.
Method
The authors build the Oytser corpus and Kashes benchmark, then continue pretraining Llama 3.1 8B to create MameLoshnLM.
Results
62.6 average score across 14 evaluations, ahead of Gemma-2 9B (57.0), Llama 3.1 8B (56.8), and Qwen3 8B (54.7), with gains concentrated on Yiddish-centered tasks.
Takeaways & Limitations
The work provides a foundation for Yiddish NLP and a useful reference point for developing models for other underrepresented languages.
Takeaways & Limitations
The training uses web-native material under a fair-use and analogous research-exception legal basis, limiting the scope of the released model's licensing.
Abstract
from arXiv · showhide
We present MameLoshnLM, the first open-source 8B-parameter language model built specifically for Yiddish. Despite Yiddish's rich textual tradition, its limited digital presence and the scarcity of reliable evaluation resources have constrained progress in Yiddish language modeling. Existing multilingual corpora and benchmarks are often poor proxies for the language, containing substantial amounts of noisy, machine-translated, and misclassified text. We address these gaps by introducing Oytser, a high-quality Yiddish pretraining corpus that combines contemporary web-native sources with literary materials, and Kashes, a multi-task benchmark spanning translation, linguistic analysis, information extraction, and language understanding. Using these resources, we continue pretraining Llama 3.1 8B to obtain MameLoshnLM. Across the tasks in the benchmark, MameLoshnLM outperforms open baselines of similar scale. Our analyses show that these gains are not only quantitative: relative to general-purpose multilingual models, MameLoshnLM better captures language-defining lexical and morphological patterns, pointing to a broader failure mode of noisy web-scale multilingual data for low-resource languages. Our results provide both a foundation for Yiddish NLP and a practical template for language model development in historically rich but digitally underrepresented languages.
1 Introduction
Yiddish language modeling is constrained by scarce, noisy training data and inadequate evaluation resources. MameLoshnLM addresses these gaps through targeted corpus and benchmark construction, achieving stronger Yiddish performance and more natural lexical and morphological behavior than comparable general-purpose models.
- Motivation: Yiddish remains underserved because multilingual language-model performance is strongest for languages with extensive digital presence, high-quality text, and mature evaluation resources.
- Motivation: Much putative Yiddish data in mC4 is machine-translated spam or not Yiddish, while existing benchmarks are scarce and often translation-based rather than language-specific.
- Contributions: The authors introduce MAMELOSHNLM, an open-source 8B-parameter Yiddish language model, alongside Oytser, a corpus combining web-native and literary materials, and Kashes, a Yiddish evaluation benchmark.
- Results: MAMELOSHNLM is obtained by continuing pretraining Llama 3.1 8B and outperforms open models of similar scale on translation, linguistic analysis, and named entity recognition.
- Broader significance: The results indicate that authentic-language mismatch in noisy public web data, not only data scarcity, can hinder modeling historically rich but digitally underrepresented languages.
2 Related Work
Continued pretraining of strong open-weight models has adapted language models to several low- to moderate-resource languages, while questions remain about training data composition, language mixing, and machine translation. Yiddish NLP has foundational prior work but no prior effort targeting modern language-model development; this work addresses that gap with a dedicated corpus, benchmark, and open-source model.
- Continued pretraining of strong open-weight models has proven effective for adapting language models to low- to moderate-resource languages, including Basque, Estonian, Kazakh, and Setswana.
- This adaptation paradigm leaves open questions about training data composition, the role of language mixing, and the value of machine translation.
- Although Yiddish has received limited NLP attention, prior efforts established foundations through phrase-based machine translation, dependency annotation, and domain-specific named-entity recognition.
- No prior work, to the authors’ knowledge, has targeted Yiddish through the lens of modern language-model development.
- The work addresses this gap by combining a dedicated large-scale pretraining corpus, a broad Yiddish language-model benchmark, and an open-source Yiddish model.
3 Training Data
The section shows that Yiddish data in mC4 is substantially noisy and introduces Oytser as a verified alternative combining contemporary web-native and literary sources. Oytser contains more than 915M words from nine high-quality online and literary sources, including the Yiddish Book Center Corpus.
- mC4 analysis: Only 42.2% of documents in mC4’s Yiddish split are genuine high-quality Yiddish from validated native sources.At least 29.8% appear machine-translated, 21.9% are Hebrew texts misidentified as Yiddish, and 6% comprise fragments, multilingual pages, and small uncatalogued sources.
- Oytser corpus: Oytser was constructed as a higher-quality alternative to automatically assembled multilingual corpora, combining contemporary casual texts with literary materials across domains and registers.The corpus was designed to reflect genuine Yiddish usage and its richer textual tradition rather than relying only on limited contemporary web presence.
- Web-native sources: More than 0.3M native Yiddish documents in Oytser come from news, magazines, Wikipedia, a Hebrew Bible translation, and freer-style forums spanning topics, dialects, and communities.These web-native sources provide contemporary and diverse forms of Yiddish, including less standard user-generated writing.
- Yiddish Book Center Corpus: More than 12K books and approximately 720M Yiddish words come from the Yiddish Book Center’s digital library, described as the largest and most comprehensive digital source currently available.The books were OCRed using Jochre3 and provide access to Yiddish’s long literary history.
- Oytser corpus: More than 915M words from 9 high-quality online and literary sources comprise the completed Oytser corpus.The resources together cover multiple genres and registers.
4 Kashes: The Yiddish Evaluation Benchmark
Kashes is introduced as the first multi-task benchmark for Yiddish language models, addressing the scarcity of systematic evaluation resources across linguistic analysis, information extraction, translation, and language understanding. It also includes Kashes-mt, a native-authored Yiddish–English translation corpus constructed from bilingual publications through alignment and quality filtering.
- Motivation: Kashes addresses the scarcity of Yiddish evaluation data, for which prior work had not systematically benchmarked language models across multiple tasks.Available Yiddish evaluation resources are described as few and scattered.
- Benchmark scope: Kashes is the first multi-task evaluation benchmark for Yiddish language models, spanning nine tasks across linguistic analysis, information extraction, translation, commonsense reasoning, and paraphrase detection.The benchmark consolidates existing resources and contributes a new parallel corpus for machine translation.
- Language understanding: Because native Yiddish benchmarks for general language understanding are unavailable, Kashes uses selected machine-translated Aya tasks whose translations preserved the essential meaning.These tasks are PIQA, WikiQA, and PAWS-Wiki.
- Kashes-mt: Kashes-mt uses bilingual online publications whose Yiddish texts were originally written by native speakers and whose English translations were produced by scholars and professional translators.The sources are Forverts, a digital newspaper, and In geveb, a peer-reviewed Yiddish studies journal.
- Kashes-mt: Kashes-mt applies document matching, sentence alignment, quality filtering, and deduplication to construct sentence-level parallel data.Documents with fewer than 35% aligned sentences and sentence pairs with SentAlign scores below 0.65 are discarded.
5 MAMELOSHNLM
MAMELOSHNLM is presented as the first open-source large language model for Yiddish, created by continued pretraining of Llama-3.1-8B on a Yiddish corpus. It outperforms similarly scaled strong baselines across a broad set of Yiddish evaluation benchmarks.
- MAMELOSHNLM is the first open-source large language model developed specifically for Yiddish.The name refers to the traditional Yiddish term meaning “mother tongue.”
- MAMELOSHNLM outperforms strong baselines of similar scale across a broad set of Yiddish evaluation benchmarks.
- Training Details: The model was produced by continued pretraining Llama-3.1-8B on a Yiddish corpus using a causal language modeling objective.Training used bfloat16 precision, 8-bit AdamW, a 2 × 10−5 learning rate, cosine scheduling, and a 2% warmup ratio.
- Training Details: Training used approximately 5.7 billion tokens over 36,663 optimization steps on a single NVIDIA H200 GPU.The data comprised 72% of words from Yiddish and 28% from English, with multilingual variants also tested using German, Hebrew, Polish, and Russian.
6 Experimental Setup
The evaluation compares MameLoshnLM with five open-weight models of similar scale. Qwen3 8B is the only baseline explicitly documented as supporting Yiddish.
- Baselines: MameLoshnLM is evaluated against Llama 3.1 8B, Qwen3 8B, BLOOMZ 7B, Gemma-2 9B, and EuroLLM 9B, with Qwen3 the only baseline explicitly listing Yiddish support.The remaining models do not claim explicit Yiddish support but serve as comparable-size open-weight baselines.
7 Results
MAMELOSHNLM achieves the strongest overall results across 14 evaluations, with particularly large gains on Yiddish-centered tasks and English-to-Yiddish translation. It remains competitive on general benchmarks while underperforming mainly on machine-translated Aya benchmarks.
- Overall performance: 62.6 average score makes MAMELOSHNLM the best overall model across 14 evaluations, ahead of Gemma-2 9B (57.0), Llama 3.1 8B (56.8), and Qwen3 8B (54.7).It leads on POS tagging, dependency parsing, transliteration, EHRI NER, and newNLP NER.
- General benchmarks: MAMELOSHNLM remains competitive on WikiANN NER and WikiQA, while its main exceptions are PAWS-Wiki and PIQA.PAWS-Wiki and PIQA are Aya benchmarks based on machine-translated versions of widely used multilingual datasets.
- Translation: MAMELOSHNLM improves English-to-Yiddish translation by more than 11 COMET points over the closest competitor.The authors attribute this strength to continued training enabling fluent Yiddish generation.
- Interpretation: Continued pretraining on authentic Yiddish data substantially improves tasks requiring lexical, orthographic, and syntactic command while preserving competitive performance on general benchmarks.This overall pattern supports the value of authentic Yiddish data for language-specific capabilities.
8 Analysis
Analysis shows that general multilingual models can generate understandable Yiddish while systematically underrepresenting its loshn-koydesh vocabulary and Yiddish-specific morphology. MameLoshnLM more closely matches native-Yiddish patterns, especially where lexical or inflectional forms are not recoverable through simple surface copying.
- Loshn-koydesh vocabulary: MameLoshnLM produces 4.7% loshn-koydesh content tokens, compared with 1.6% for Llama 3.1 8B, versus 6.2% in gold references.The comparison covers 5,287 English→Yiddish 5-shot translations; p < 10^-229 by paired t-test.
- Loshn-koydesh vocabulary: Llama often replaces loshn-koydesh nouns with Germanic alternatives and avoids function words through paraphrase or restructuring, yielding more Germanized Yiddish.The passage gives examples involving family, war, face, perhaps, almost, and in order to.
- Yiddish-specific morphology: For citation-lemma accuracy on דָאָס, די, and דעם forms corresponding to English “the,” MameLoshnLM reaches 28.2% versus Llama’s 15.5%.These surface forms must map to the same citation lemma, דער.
- Yiddish-specific morphology: On regular האָבן forms, both models reach 84.7%, but on irregular זײַן forms MameLoshnLM reaches 20.0% versus Llama’s 8.9%.The contrast suggests both models handle locally recoverable lemmas, while MameLoshnLM is stronger when inflected forms differ substantially from the lemma.
- Data quality and model behavior: Machine-translated Yiddish pages contain 3.6% loshn-koydesh tokens, compared with 10.2% in validated native sources, indicating weaker native-language signal in noisy web corpora.The passage presents this audit as evidence that machine-translated material contributes to the lexical deficit observed in general multilingual models.
9 Discussion and Conclusion
The discussion frames Yiddish as both a practically important language requiring dedicated technology and an informative testbed for multilingual NLP. The work establishes a comprehensive Yiddish LLM development foundation while identifying directions for future research and digital humanities applications.
- Motivation: Yiddish’s roughly one million speakers and vast textual heritage motivate language technologies for searching, organizing, and analyzing that heritage at scale.The passage notes that inadequate language technologies make Yiddish’s textual heritage difficult to handle at scale.
- Yiddish as a multilingual NLP testbed: Yiddish’s position across language families and its unique script make it a testbed for tokenization, cross-lingual transfer, multilingual data mixing, and data scarcity.The passage presents Yiddish as informative for studying how data scarcity and quality affect language model development.
- Contributions and future work: The work constitutes the first comprehensive Yiddish LLM development effort, spanning corpus construction, benchmark curation, model training, and evaluation.Future work includes instruction tuning, additional training and evaluation resources, and applying MAMELOSHNLM to large-scale digital humanities workflows.
Ethics Statement … C Kashes Examples
The paper documents its institutional, copyright, and privacy considerations while specifying alignment filters for Kashes-mt and presenting few-shot evaluation results and task examples. These materials describe how the corpus and benchmark were assembled, filtered, and formatted.
- Ethics Statement: Yiddish Book Center and In geveb partnerships supported model training, evaluation, preservation, accessibility, and research.The Yiddish Book Center licensed the Steven Spielberg Digital Yiddish Library, while In geveb licensed evaluation material.
- Ethics Statement: More than 82% of training tokens came from institutional agreements, open or public-domain licenses, and applicable source terms.Web-native text was collected for non-commercial academic research from pages accessible without login, subscription, paywall, or circumvention.
- Ethics Statement: The training-stage use of publicly available web data is framed as non-expressive, transformative computational research relying on fair use and analogous research exceptions.The stated purpose is to learn general linguistic patterns rather than provide access to or substitute for individual works.
- Ethics Statement: Ivelt and Kaveshtiebel were included as substantial public contemporary Yiddish sources because they represent everyday community language largely absent from historical and edited collections.The passage also notes that both sources were already prominent in Common Crawl-derived training data.
- A Kashes-mt Sentence Alignment: Documents with less than 35% initially aligned sentences were removed, and sentence pairs with SentAlign similarity scores below 0.65 were filtered as potential inaccurate translations.Figure 2 shows document-level alignment shares, while Figure 3 shows similarity-score distributions.
- B Few-shot full result: Table 6 reports results across all tasks and shot counts, with boldface indicating the best result per row.The supplied passage provides the table’s scope and formatting convention but no individual metric values.
- C Kashes Examples: Tables 7–10 provide few-shot prompt-format examples for Aya Collection, UD Yiddish-YiTB, NER, and machine translation tasks.Quoted English text beneath examples is supplied only for reader translation and explanation, not as evaluation input.
D Auditing the Yiddish Split of mC4 … E.7 Statistical Testing
The appendix audits substantial contamination in mC4’s Yiddish split and documents the lexicon-, morphology-, and statistical-testing procedures used to evaluate Yiddish linguistic competence. It combines targeted diagnostics with paired comparisons between MameLoshnLM and Llama 3.1 8B.
- D Auditing the Yiddish Split of mC4: 6.2% of the split forms a long tail of 8,871 difficult pages, so both native-Yiddish and machine-translation shares are lower bounds.The residual pages include short Hebrew-script fragments, multilingual templates, and small uncatalogued native sources.
- E Analysis Methodology; E.1 Loshn-Koydesh Identification: The appendix’s linguistic analyses use a loshn-koydesh lexicon and UD Yiddish-YITB morphological annotations, with the lexicon containing 5,437 single-word entries and 2,926 compounds.LK matching uses NFC-normalized exact matches, diacritic-stripped fallback, and sliding-window compound matching.
- E.2 LK Vocabulary Production in Translation: 5,287 English–Yiddish sentence pairs support 5-shot LK translation analysis, measuring LK content-word rate, sentence match rate, and per-word LK recall.Substitution analysis examines 72 frequent gold LK words and identifies empirically observed Germanic equivalents.
- E.3 Morphological Analysis via Lemmatization: 929 sentences containing 7,499 tokens form the lemmatization test set, with models evaluated on structured word-to-lemma predictions and LK versus non-LK change accuracy.The dataset contains 565 LK tokens (7.5%) and 6,934 non-LK tokens (92.5%); 48.1% of LK tokens require lemma changes versus 33.5% of non-LK tokens.
- E.4 Morphological Category Breakdown: 2,593 change tokens are classified into six morphology categories in priority order, assigning each token exactly one category for category-level change-accuracy analysis.Categories use gold CoNLL-U UPOS tags to disambiguate phenomena such as ge- participles, adjective declension, and Hebrew-origin plurals.
- E.5 Determiner Case System; E.6 Auxiliary Verb Paradigms: The determiner analysis groups gold-lemma דער tokens by surface form and excludes the identity form, while the auxiliary analysis compares 460 suppletive זײַן tokens with 163 regular האָבן tokens.The auxiliary comparison tests whether language-specific knowledge benefits suppletive morphology more than straightforward suffix changes.
- E.7 Statistical Testing: McNemar’s exact test evaluates paired binary model outcomes, while paired t-tests assess continuous per-sentence LK rate and recall in translation.Discordant outcomes define b and c, with significance reported at p < 0.05, p < 0.01, and p < 0.001.
E.8 Summary of All Metrics · E.9 Most Frequent Loshn-Koydesh Words in Gold References · F Effect of related-language mixing
The paper consolidates MAMELOSHNLM–Llama 3.1 8B comparisons across translation and morphological analyses, examines Loshn-Koydesh word frequencies in gold references, and evaluates mixtures with more Yiddish-related languages. These analyses cover 5-shot settings and compare outputs, metrics, and training-mixture compositions.
- E.8 Summary of All Metrics: Table 12 consolidates translation and morphological metrics comparing MAMELOSHNLM with Llama 3.1 8B.Translation uses 5-shot evaluation, while morphology uses 5-shot lemmatization.
- E.8 Summary of All Metrics: Translation metrics use 5,287 sentence pairs, with match and recall computed over 3,112 gold-LK sentences.The comparison is conducted in a 5-shot setting.
- E.8 Summary of All Metrics: Morphological metrics are evaluated 5-shot on the UD Yiddish-YITB test set, with McNemar b:c denoting discordant pairs favoring MAMELOSHNLM.All p-values in the consolidated comparison are two-sided.
- E.9 Most Frequent Loshn-Koydesh Words in Gold References: Table 13 reports the 15 most frequent Loshn-Koydesh content words in gold Yiddish references and their absolute output counts.Counts cover 5-shot translations across 5,287 sentence pairs from the In Geveb and Forward sources.
- E.9 Most Frequent Loshn-Koydesh Words in Gold References: Common Yiddish vocabulary such as efsher, ponim, and kedey is largely absent from Llama’s output, whereas rebe, khane, and yisroel appear at comparable rates.The comparison concerns absolute counts in the 5-shot translation outputs.
- F Effect of related-language mixing: Table 14 compares MAMELOSHNLM and Llama 3.1 with models trained using mixtures containing more Yiddish-related languages.All Kashes results use a 5-shot setting and are scaled from 0 to 100, with bold indicating the best per row.
- F Effect of related-language mixing: Table 15 gives the training-mixture composition for three continued-pretraining settings, reporting percentages over words and training tokens.The table describes the mixture configurations underlying the related-language comparison.