Source-linked AI summary

Language, Language Models, and What We're Talking About

Malvina Nissim

arXiv:2609.03577v1cs.CL

TL;DR

The paper asks whether models trained and evaluated through translated, synthetic, and highly curated data should be understood as models of Italian or of language. It examines Italian models, their evaluation practices, and alignment, arguing that technical products and tools for studying language require distinct goals, data relationships, and evaluation criteria.

  • Problem

    Current NLP often uses the same modelling and evaluation strategies for building language technologies and studying language, despite their different requirements for data and linguistic fidelity.

  • Method

    The paper uses Italian language models as a case study, reviewing their construction and evaluation while reflecting on alignment and the meaning of language modelling.

  • Results

    Italian models and evaluations combine highly curated, English-based systems with translated or synthetic Italian data and artificial benchmarks that may say little about language production.

  • Takeaways & Limitations

    Language-model research should distinguish technical products from tools for studying language, with separate goals, methodologies, and evaluation criteria.

  • Takeaways & Limitations

    The proposed distinction is not yet a concrete, well-defined proposal and is not presented as a settled position.

Abstract

from arXiv · show

Language models are commonly discussed as technical artefacts, but they are obviously shaped by the linguistic worlds conveyed by data during their training. Using Italian language models as evidence, I want to bring attention to the nature of the systems which result from training and specialising models on translated and synthetic data, and further curating them, and to the meaning of testing them on equally unnatural data. Are these eventually models of Italian? Are they models of language? Does NLP still care about language? These questions yield another, more concrete question: what language do we actually want language models to produce? I argue that this question cannot be answered if we do not first consider a clearer distinction between language models designed as technical products and language models designed as tools for studying language itself. The answers then might be diverse, the languages we are talking about might be diverse, and the picture might not be as pessimistic as we fear.

1. The State of Things

LLMs are technical artefacts whose behaviour emerges from linguistic input across training and refinement stages. The paper uses Italian models to question whether current practices still meaningfully concern language.

  • 1. The State of Things: An LLM is a technical artefact whose behaviour is shaped by the linguistic world conveyed through its training data.That world emerges from the linguistic input processed during model creation and refinement.
  • 1. The State of Things: Model development typically involves pre-training, instruction-tuning, and alignment-driven adjustment based on human preferences and policy optimisation.These stages form the paper’s simplified reference structure, despite later extensions such as tool access and inference-time reasoning.
  • 1. The State of Things: Most influential LLMs are primarily trained on English, while non-English specialisation often relies on limited multilingual exposure or later adaptation.The paper contrasts this pattern with models trained substantially from scratch on native data, including Minerva for Italian.
  • 1. The State of Things: Using Italian as a case study, the paper examines language in existing Italian LLMs, their evaluation, alignment effects, and what it means to model language.It extends these questions to the outputs expected from models and the future direction of NLP.

2. A brief history of Italian generative models

Italian generative models evolved from small, Italian-only systems toward larger models adapted from existing backbones with translated or synthetic data and inference-time controls. The section surveys these divergent strategies and their resulting model landscape.

  • 2. A brief history of Italian generative models: GePpeTto was the first decoder-only Italian language model trained fully from scratch on Italian-only data.It used a GPT-2-small architecture with about 14 gigabytes of Italian text and 117M parameters.
  • 2. A brief history of Italian generative models: IT5 used an Italian C4 corpus of 215GB of cleaned web text for multitask training in translation, summarisation, question answering, and style transfer.It was evaluated on ItaGen tasks covering summarisation, headline generation, and formality manipulation.
  • 2. A brief history of Italian generative models: Camoscio, Fauno, and Anita adapted LLaMA models using Alpaca-derived instruction data, including automatically translated Italian examples.Camoscio fine-tuned LLaMA 7B with LoRA on an Alpaca dataset translated into Italian using GPT-3.5.
  • 2. A brief history of Italian generative models: Steered-ITA adapted an unchanged base model to Italian at inference time through contrastive activation steering and Italian in-context examples.This provided a lighter alternative to standard instruction tuning, which updates model weights using supervised data.
  • 2. A brief history of Italian generative models: Minerva adopted a from-scratch strategy with a balanced Italian-English corpus, code, and 2.5 trillion total tokens.Its pre-training data included 1.14 trillion Italian tokens, 1.14 trillion English tokens, and 200 billion code tokens.
  • 2. A brief history of Italian generative models: The resulting landscape shifted from small Italian-from-scratch models toward larger systems reusing pre-trained backbones, synthetic or translated instructions, and inference-time controls.The paper presents these approaches as the dominant recent pattern in Italian model development.

3. Italian models my ass!

The section questions whether models built from English-centered bases and translated, synthetic, or machine-generated Italian data should straightforwardly be called Italian models. It examines their training inputs and evaluations, arguing that these choices complicate what the models represent and how their Italian abilities should be judged.

  • Base models (Stage 1): Models such as Camoscio, Fauno, Anita, and Steered-ITA originate from existing pretrained bases, including LLaMA and Phi-3, rather than being trained from scratch on Italian data.Steered-ITA adapts its base model at inference time, while the other models use instruction-tuning or related specialization procedures.
  • Instruction-tuning (Stage 2): Camoscio is an automatically translated version of Alpaca, so its Italian instruction data inherits translation artifacts from an English synthetic dataset.The dataset remains usable in many cases, but examples include cultural mismatches, lost wordplay, nonsensical grammatical analyses, and broken instruction-response relations.
  • Instruction-tuning (Stage 2): Fauno and Anita also draw on synthetic dialogues, online discussion threads, and other synthetic or machine-translated instructions, while ALERT is described as native Italian instruction-following data.The section presents these sources as differing kinds of instruction-tuning data rather than as a single uniform Italian corpus.
  • Evaluation: Italian models are commonly evaluated on downstream tasks including summarisation, question answering, formality transfer, and translated or Italian-specific benchmarks.The evaluation landscape includes the Open ITA LLM Leaderboard, native Italian datasets, and translated versions of English tasks such as MMLU.
  • Conclusion: The section asks whether models built on English-pretrained systems and adapted and evaluated with translated or synthetic data are meaningfully Italian models.This question is presented as extending beyond Italian to the broader status of language in contemporary model training and evaluation.

4. What do you mean by Language Model?

The section contrasts formal and usage-based conceptions of language before arguing that contemporary NLP may have lost interest in natural language itself. It questions evaluations that reward correct answers while treating artificial or homogenised output as if it were natural language.

  • From formal systems to usage: Linguistics shifted from formal, top-down accounts of language toward ecological and usage-based approaches focused on situated use.The section traces this movement through structuralist and generative views toward approaches that model language as it is actually used.
  • NLP and language: The author argues that current NLP shows limited interest in natural language itself, despite extensive work on multilingualism and low-resource languages.The criticism distinguishes concern for language from concern for linguistic theories and their role in model development.
  • Evaluation and output: Artificial benchmarks often prioritize answer correctness over language production, rewarding numbers, formal expressions, or option letters rather than natural textual output.A correct result therefore may provide little information about how a model produces language.
  • What should models produce?: The section questions whether homogenised, fluent, deliberately artificial language should be judged by naturalness, and asks what language models should actually produce.It frames the desired character of model output as an unresolved question rather than assuming that naturalness is the proper target.

5. Some Optimism and Many More Open Questions

The paper argues that language-focused and application-focused models require distinct goals, data practices, and evaluation criteria. Separating these directions could clarify how natural and artificial language should relate in model development and use.

  • Native benchmarks and pluralistic alignment recognise important aspects of language and use, but their role remains unclear for highly curated or artificial models.The paper specifically questions the meaning of testing English-based models specialised with synthetic Italian on native Italian benchmarks.
  • Contemporary NLP often pursues language study and application development through the same modelling and evaluation strategies, despite their different relationships with data.Linguistic and cultural study prioritises ecological fidelity, while applications require curation and alignment as social interventions.
  • Combining authenticity and intervention produces contradictory outcomes by seeking to preserve linguistic diversity while standardising the variation models could represent.The paper argues that these objectives cannot be treated as compatible optimisation goals.
  • Applied models: An applied direction could use heavily curated models that produce recognisable artificial language rather than pretending to produce natural language.Such systems would be evaluated as assistive tools, with explicit decisions about bias, values, and the artificial languages they represent.
  • Language-research models: A language-research direction could use smaller models trained on ecological data, minimally altered and intended to support linguistic and cultural inquiry rather than deployment.These models may be commercially unattractive but potentially more useful for studying language and culture.
  • The proposed distinction is not yet a concrete, well-defined, or fully endorsed proposal, but is offered as a starting point for discussion.The author presents it as a more articulated reflection rather than a settled framework.

Declaration on Generative AI

The author used Claude and Nano Banana for image generation and ChatGPT for grammar checking and selected rephrasing, then reviewed and edited the content.

  • Claude and Nano Banana were used to generate images.
  • ChatGPT was used for grammar checking and some rephrasing where style was not a priority.
  • The author reviewed and edited the generated contributions and takes full responsibility for the publication’s content.
Loading 2609.03577v1…