Source-linked AI summary
Language Models for Portuguese: A Systematic Mapping Study
Jhessica Silva, Carlos Caetano, Helena Maia, Breno Bernard Nicolau de França, Sandra Avila, Helio Pedrini
TL;DR
Information about Portuguese language models is fragmented across sources, limiting comprehensive characterization of the field. This systematic mapping study analyzes 46 models across multiple dimensions and identifies representation gaps and future research opportunities.
Problem
Portuguese language-model information is fragmented across publications, repositories, and documentation, leaving key aspects insufficiently characterized.
Method
The study systematically maps Portuguese language models, characterizing their architectures, resources, datasets, availability, evolution, relationships, and research gaps.
Results
The mapping identified 46 Portuguese language models, with most published from 2023 onward and development concentrated in Brazil and Portugal.
Takeaways & Limitations
Future work should develop models tailored to Portuguese variants, expand multimodal coverage, conduct qualitative and ethical analyses, and improve model sharing.
Takeaways & Limitations
The study covers Brazilian Portuguese, European Portuguese, or both, which does not represent the full diversity of Portuguese-speaking populations.
Abstract
from arXiv · showhide
In recent years, the rapid development of language models has transformed the field of Natural Language Processing through a wide range of applications. However, the development of language models has not progressed uniformly across all languages. In the case of the Portuguese language, there has recently been a growing effort by academia and companies to develop language models and create data resources for Portuguese. These efforts have resulted in the rise of an increasingly diverse ecosystem of language models for Portuguese. However, information on these models remains dispersed in scientific publications, technical reports, model repositories, and project documentation. This survey presents a systematic mapping study of language models developed for Portuguese, providing a comprehensive overview of the current state of the field. We map a total of 46 models, characterizing them by various aspects, including base model, architecture, computational resources, training datasets, licensing, code availability, data, and model weights. Furthermore, we analyzed the evolution and relationships among these models through a phylogenetic perspective, identified current research gaps and opportunities, and discussed future directions for the development of language models for Portuguese.
1. Introduction
Portuguese language-model development has grown but remains fragmented and constrained by unequal resource availability. This study addresses the fragmentation through a systematic mapping of Portuguese language models, characterizing their properties and analyzing their evolution, gaps, and future directions.
- Motivation: Research on language models has favored high-resource languages, while limited corpora, computational power, and benchmarks make development for other languages more challenging.Portuguese has traditionally been considered less resourced than English.
- Motivation: Recent efforts from academia and industry have expanded Portuguese language models and linguistic resources, alongside governmental initiatives recognizing their strategic importance.Examples include Brazil’s PBIA and Portugal’s ANIA.
- Research problem: Information about Portuguese language models remains dispersed across publications, reports, repositories, and documentation, hindering a comprehensive view of the field.This fragmentation also complicates assessment of availability, training data, openness, documentation quality, and architectural evolution.
- Study contribution: The study conducts a systematic mapping of Portuguese language models published between 2020 and 2025 using rigorous review and snowballing methods.The mapping is intended to provide a comprehensive overview of the field’s current state.
- Study contribution: The mapping characterizes models by base model, architecture, resources, datasets, licensing, code, data, and weights, while examining phylogenetic relationships, research gaps, opportunities, and future directions.These dimensions support analysis of how Portuguese language models have evolved and relate to one another.
2. Related Work
Prior work has surveyed Portuguese language models, evaluated Brazilian Portuguese generation, and compared models by chronology, architecture, parameters, and computational resources. This study builds on these efforts while providing its own systematic mapping of models for Portuguese.
- Surveys of Portuguese Models: 24 models released between 2020 and October 2024 were surveyed, distinguishing Brazilian Portuguese from European Portuguese models.The survey accompanied work proposing Tucano and later Tucano 2.
- Evaluation and Timelines: Six Brazilian Portuguese large language models were analyzed for generative performance on Natural Language Generation tasks.Model selection used the Open Portuguese LLM Leaderboard, presenting a timeline of 29 models released between 2020 and 2024.
- Comparative Surveys: 45 Brazilian Portuguese models released between 2020 and March 2025 were compared by architecture, parameter count, and computational resources such as energy consumption.The survey included 16 models not covered by this study because its methodology does not consider unpublished models.
3. Research Method
The study uses a systematic mapping approach to characterize Portuguese-focused language models released from 2020 onward, including unimodal and multimodal models. It combines database searching, snowballing, reviewer screening, information extraction, and quality assessment.
- Study goal: The study targets language models developed on Portuguese textual data and released from 2020 onward, including unimodal and multimodal models.The models may perform any activity as long as they focus on Portuguese.
- Research questions: The mapping study addresses which Portuguese language models exist, their model, code, and data availability, their training benchmarks or datasets, and their data origins.The research questions also distinguish Portuguese-origin data from data translated from another language.
- Search strategy: The search covered Scopus, IEEEXplore, Web of Science, and arXiv, complemented by forward and backward snowballing.The arXiv search was split into four separate queries because of platform limitations, and searches were restricted to papers published since 2020.
- Study selection: Two reviewers screened all papers using inclusion and exclusion criteria, categorizing each as Include, Exclude, or Uncertain.Papers classified as A were included, while those classified as B, C, and D received full-text eligibility assessment; E papers were excluded.
- Data extraction: The reviewers extracted model information from papers, technical reports, model cards, and repository README files, then reviewed all extracted information.The extraction form was designed to answer the study’s research questions.
- Quality assessment: Each included model was assessed using criteria covering the study, model development, and model analysis, with scores of 0 or 0.5 assigned when information was absent or partial.The supplied passage indicates that the quality criteria used three categories and included a score of 0 for unavailable information and 0.5 for partial information.
4. Review and Findings
The systematic mapping identified and characterized 46 language models for Portuguese through automated search and snowballing. The mapped models span diverse publication periods, venues, developer affiliations, Portuguese variants, purposes, base models, parameter counts, licenses, and training datasets.
- Search and selection: 32 language models for Portuguese were included after the automated search, with all control papers returned.The search was conducted on June 3, 2025, across the specified libraries after duplicate and non-paper removal.
- Search and selection: The 32 papers from automated search were used in forward and backward snowballing, with selection criteria applied by two reviewers.Snowballing was conducted on August 31, 2025, using titles of reference and citing papers.
- Publication trends: Most included models were published from 2023 onward, with the publication period spanning 2020 to August 2025 and covering Brazilian, European, and general Portuguese variants.The publication date refers to the paper or the first version on arXiv or HuggingFace; some models preceded the reported publication date.
- Publication venues: BRACIS, EPIA, and PROPOR were the most frequent publication venues, with 6, 3, and 3 publications, respectively.SIGUL and CBMS each had two publications, while venues otherwise covered diverse AI, NLP, under-resourced-language, and medical-computing topics.
- Developer affiliations: 33 universities and 14 companies were affiliated with model development, led by USP with 7 models, UNICAMP with 6, and Maritaca AI with 5.More than one institution could collaborate on a single model.
- Mapping scope: 46 language models for Portuguese were identified and characterized by publication/release year, base model, purpose, parameter count, license, and training dataset.The study presents these characteristics in Table 5.
5. Analysis and Discussion
The analysis assesses study quality, model purposes and lineage, resource availability, and training-data diversity across Portuguese language models. It identifies broad variation in model applications, foundations, access conditions, and dataset sources.
- Quality assessment: 7.5 out of 11 was the average quality score across studies, which were evaluated on 11 criteria without excluding any study.Individual scores ranged from 1.5 to 10.5 in 0.5-point intervals.
- Purposes and phylogeny: 46 identified models included 27 trained for general Portuguese text tasks and 2 for tasks combining Portuguese text and images.The models were diverse in their purposes and approaches despite all focusing on Portuguese.
- Purposes and phylogeny: 20 Portuguese models descended from BERT, forming the largest identified phylogeny, while the Llama family formed the second-largest with 10 models.The mapped models also included derivation branches such as BERT → RoBERTa → DeBERTa → Albertina-PT-* → MediAlbertina-PT*.
- Availability: Model, code, training-data, and documentation availability were evaluated using open, semi-available, and unavailable categories.Code and models counted as available when openly accessible through repositories such as GitHub or HuggingFace.
- Datasets and benchmarks: Web / Crawled Data was the largest dataset category, while Legal / Governmental datasets represented a significant share of Portuguese model training data.Clinical / Healthcare, Instruction / Conversational, and Multilingual / Para… categories were also present, indicating heterogeneous data sources.
- Datasets and benchmarks: General-purpose models predominantly used Web / Crawled Data, whereas domain-specific models typically used curated datasets tailored to their respective domains.Instruction / Conversational datasets were increasingly adopted by recent models for interactive and task-oriented applications.
6. Gaps and Opportunities
The study identifies gaps in representation, data quality, ethical and qualitative evaluation, accessibility, and multimodal modeling for Portuguese. It highlights opportunities to develop culturally grounded resources, transparent documentation, accessible models, and more diverse multimodal systems.
- The Portuguese language and diversity: Portuguese language models largely target Brazilian Portuguese, European Portuguese, or both, overlooking the language’s broader pluricentric diversity.Portuguese is official in nine countries and present across four continents, so models should address variant-specific cultural, historical, social, and economic characteristics.
- Translated data and data quality: Machine-translated English-to-Portuguese data can reproduce global-North perspectives and translation errors, while newer Portuguese corpora create opportunities for native data resources.The mapping reports that new Portuguese training corpora and datasets were created between 2020 and 2025, particularly for Brazilian Portuguese.
- Qualitative and ethical analysis: Only 5 of 46 mapped models discussed ethical issues, 17 addressed limitations, and 8 conducted qualitative analysis.Although all models included quantitative evaluation using literature-defined metrics, ethical and qualitative analyses remained uncommon.
- Qualitative and ethical analysis: Model documentation should cover intended uses, target populations, underrepresented groups, and social, cultural, economic, and environmental impacts, including development-cycle carbon footprints.Existing carbon-footprint reporting often measured only the final model version without a robust method.
- Model availability and access methods: Only 11 of 46 mapped models had source code available, limiting transparency, reproducibility, and progress, while access through repositories still favors technically skilled users.Most models were available for community use through HuggingFace or GitHub, but the passage notes that practical access remains constrained by technical expertise.
- Multimodal models: The search identified only two vision-language models, motivating research on additional modalities, advanced multimodal backbones, and regionally or culturally grounded images.Existing studies generally use English-centered image datasets paired with translated text, leaving regional visual representation as an open opportunity.
7. Conclusion
This systematic mapping study analyzes Portuguese language models published from January 2020 to August 2025, identifying 46 models and summarizing the field’s current state. It highlights linguistic, representational, ethical, and analytical gaps and proposes variant-specific, culturally grounded, ethically evaluated, and multimodal future development.
- Conclusion: 46 language models were identified through searches of Scopus, IEEEXplore, Web of Science, and arXiv, supplemented by forward and backward snowballing.Studies were filtered using inclusion and exclusion criteria.
- Conclusion: Existing models often reduce Portuguese to Brazilian and European variants, rely on culturally unrepresentative translated data, and lack qualitative and ethical analysis.The study also notes limited discussion of these issues.
- Conclusion: Future research should tailor models to each Portuguese variant, curate country-representative datasets, conduct qualitative and ethical analyses, and advance multimodal models.The authors argue that careful examination is necessary to assess whether models adequately serve a language and population.
Appendix A. Access to Language Models for Portuguese
Appendix A catalogs access to Portuguese language models as of April 2026, distinguishing models that are publicly unavailable from those linked through Hugging Face or GitHub. Most listed models have public repository or collection links, while Sabiá-3 is explicitly unavailable.
- Access overview: Sabiá-3 is not publicly available.Its entry provides no public repository link.
- Public access: Tucano, Carvalho_pt-gl, Serafim PT*, V-GlórIA, BERTugues, and BERTweet.BR have public Hugging Face or GitHub links.These entries point to model collections or repositories hosted on Hugging Face or GitHub.
- Public access: GovBERT-BR, PTT5-v2, and Clinical-BR-* also have Hugging Face links, whereas Amadeus-Verbo is listed with a Hugging Face collection link.The Clinical-BR-* entry includes two model links, and Amadeus-Verbo is paired with a separate collection URL.
Appendix B. Authors’ Affiliations Acronyms List
Appendix B decodes acronyms for authors’ affiliated organizations, spanning companies, research groups, and universities in Brazil and abroad.
- Affiliation acronyms: The list expands acronyms for organizations including 22h, Alfaneo, Amadeus AI, BotBot, Comsentimento, Datalab, FEI, HES-SO, PUC-Rio, PUC-RS, Sagui AI, and Select Data.It also identifies Datalab as Latam Datalab Serasa Experian, FEI as Centro Universitário da Fundação Educacional Inaciana, and HES-SO as the University of Applied Sciences and Arts of Western Switzerland.
- Affiliation acronyms: The list also expands Tikal Tech, U.Porto, UÉ, UFF, UFG, UFMA, UFMG, UFMS, UFRGS, UFRPE, ULisboa, UMich, and UNESP.These correspond to technology organizations and universities including the Universidade do Porto, Universidade de Evora, Universidade Federal Fluminense, Universidade Federal de Goiás, Universidade Federal do Maranhão, Universidade de Lisboa, and University of Michigan.