

Multilingual AI · Data for AI
What Is a Low-Resource Language?
A language is not poor because it lacks words. It is "low-resource" because it lacks machine-readable data, tools, and benchmarks. This guide explains the definition, how the condition is measured, where it bites (from Valencian to Quechua, from Medumba to Lao), and what anyone training or buying AI for these languages should ask before trusting a model that says it "supports" them.
By Manuel Herranz, Founder and CEO, Pangeanic · Published September 27, 2026 · 20 min read
Definition
Low-resource language
A low-resource language (also called an under-resourced, resource-scarce, or low-density language) is a language for which there is not enough digital text, speech, parallel data, annotated data, software tooling, or evaluation benchmarks to build and measure language technology (machine translation, speech recognition, large language models) that performs as well as it does for English. The label describes a data condition, not the language itself, and it changes with the task, the domain, the modality, and the moment of evaluation.
Opposite term: high-resource language (English, Spanish, German, Japanese, and French sit at the top of the most cited taxonomy). Related but distinct: minority, minoritized, and endangered languages.
01 · The concept
A data condition, not a linguistic verdict
The term was born inside natural language processing (NLP) and computational linguistics, and it carries the bias of its birthplace: it measures a language by what a machine can learn from it. A language with a thousand-year literary tradition, a rich morphology, and millions of speakers can still be "low-resource" if very little of it exists as clean, licensed, machine-readable text or transcribed speech. Conversely, a language with fewer speakers but decades of public investment in corpora, spell checkers, treebanks, and translation memories may not be low-resource at all. Speaker numbers and digital resources are two different axes, and AI only sees the second one.
The most useful academic framing comes from Hedderich and colleagues (NAACL 2021), who describe a low-resource scenario rather than a low-resource language. The scenario is defined along 3 dimensions: the availability of labeled data for the specific task, the availability of unlabeled text in the language or domain, and the availability of auxiliary data (dictionaries, related languages, knowledge bases). This is why the same language can be well served for spell checking and badly served for speech recognition, or strong in news text and weak in clinical notes. Indeed, even English becomes a low-resource scenario in a narrow technical domain with no annotated data. Haddow and colleagues (Computational Linguistics, 2022) add a time dimension for machine translation: some language pairs considered scarce a decade ago now have far more data thanks to web crawling and the use of monolingual text.
Of course, the label has a cost. Some researchers point out that "low-resource" defines a language by what it lacks and can hide the knowledge, practices, and priorities of its speakers. I return to that critique later, because it has practical consequences for how data should be collected. For now, the working definition matters because procurement, research funding, and model release notes all use it, often without saying which of the 3 dimensions they mean.
Terminology
Is a low-resource language the same as a minority or endangered language?
No. The categories overlap often, but each one is defined by a different criterion. Mixing them up leads to bad policy and worse datasets.
| Term | Defined by | How it relates to AI resources |
|---|---|---|
| Low-resource language | Scarcity of machine-readable data, tools, and benchmarks for a given task | The computational criterion itself; varies by task, domain, modality, and year |
| High-resource language | Abundant labeled and unlabeled data, mature tooling and evaluation | English, Spanish, German, Japanese, and French form the top class in Joshi et al. (2020) |
| Minority language | Sociopolitical position within a territory or state | Often coincides with scarce data, but a minority language can be well resourced |
| Minoritized language | A historical process of social displacement by a dominant language | Displacement usually reduces the text produced in public domains, which later becomes training data |
| Endangered language | Interrupted or weakened intergenerational transmission | Most endangered languages are low-resource; most low-resource languages are not endangered |
02 · Inventory
What counts as a "resource"?
There is no universal threshold of tokens or hours above which a language stops being low-resource. The resources needed depend on the application, and so do the gaps. This is the inventory we use when scoping a data project.
| Resource layer | Typical resources | Frequent gaps in low-resource languages | What breaks in AI |
|---|---|---|---|
| Text | Monolingual corpora, digitized books and archives, news, encyclopedias, public administration text | Small volume, narrow domains, duplicates, broken encoding, text labeled with the wrong language | LLM pretraining, fluency, and knowledge of local context |
| Translation | Parallel corpora, translation memories, bilingual glossaries and terminology | Few sentence pairs, concentration in religious or administrative text, low-quality translations | Machine translation quality, cross-lingual transfer, terminology consistency |
| Speech | Recordings, transcriptions, pronunciation lexicons, speaker metadata | Few transcribed hours, under-represented varieties and accents, missing metadata | ASR, TTS, voice assistants, transcription of public proceedings |
| Annotation | Part-of-speech tags, named entities, dependencies, sentiment, intents, preference data | Tiny samples, incompatible guidelines, too few qualified annotators | Fine-tuning, RLHF and DPO alignment, safety and hate speech detection |
| Infrastructure and evaluation | Keyboards, fonts, tokenizers, morphological analyzers, spell checkers, benchmarks, metrics | Unstable orthographies, incomplete script coverage, restrictive licenses, no test sets | You cannot tell whether a system works, or for which variety |
The existence of a file does not make it a resource. Quality, representativeness, documentation, format, and license decide whether data can be reused. The clearest warning comes from Kreutzer and colleagues (TACL, 2022), who manually audited 205 language-specific corpora released with 5 major web-crawled datasets (CCAligned, ParaCrawl, WikiMatrix, OSCAR, and mC4). Lower-resource corpora showed systematic problems: at least 15 corpora contained no usable text at all, and a significant fraction contained less than 50% sentences of acceptable quality, with many mislabeled or tagged with ambiguous language codes. A raw token count can therefore paint an imaginary picture of coverage.
Wikipedia, the default proxy for "unlabeled text," has its own pathologies. Tatariya and colleagues (2024) found script and language contamination, placeholder articles, and heavy bot-generated content in low-resource editions, and showed that models trained on filtered data match or beat models trained on the raw dump. Cebuano is the textbook case: one of the largest Wikipedias by article count, much of it generated automatically. And there is a feedback loop worth naming. Low-quality machine translation published on the web gets crawled back into training corpora, so the language that most needs good data is the one most exposed to recycled errors. For us at Pangeanic, this is the argument for human validation by native speakers at the point where data enters a pipeline (a theme I have written about in why multilingual AI data quality is hard to get right).
03 · Measurement
How is resource level measured?
The most cited attempt to rank the world's languages by resources is "The State and Fate of Linguistic Diversity and Inclusion in the NLP World" by Joshi, Santy, Budhiraja, Bali, and Choudhury (ACL 2020). The authors built a taxonomy of 6 classes on 2 axes: the number of labeled resources (drawn from the LDC catalog and the ELRA map) and the volume of unlabeled text (approximated through Wikipedia). The result is a steep pyramid.
| Class | Name in the study | Resource profile | Example languages | Languages in class |
|---|---|---|---|---|
| 0 | The Left-Behinds | Virtually no labeled or unlabeled data | Dahalo, Warlpiri, Popoloca, Wallisian, Bora | 2,191 |
| 1 | The Scraping-Bys | Some unlabeled text, almost no labeled data | Cherokee, Fijian, Greenlandic, Bhojpuri, Navajo | 222 |
| 2 | The Hopefuls | Small amounts of both, active communities | Zulu, Konkani, Lao, Maltese, Irish | 19 |
| 3 | The Rising Stars | Strong web presence, few labeled resources | Indonesian, Ukrainian, Cebuano, Afrikaans, Hebrew | 28 |
| 4 | The Underdogs | Large unlabeled data, fewer labeled resources than class 5 | Russian, Hungarian, Vietnamese, Dutch, Korean | 18 |
| 5 | The Winners | Abundant labeled and unlabeled data | English, Spanish, German, Japanese, French | 7 |
Source: Joshi et al. (2020), ACL. Snapshot of the resource landscape at the time of publication; class membership depends on the catalogs and indicators used and is not a permanent ranking.
Look at class 1 and class 2 side by side. Bhojpuri, spoken by tens of millions of people in India and Nepal, sits in "The Scraping-Bys"; Maltese, with a speaker population well under a million, sits one class higher. Maltese is an official language of the European Union, and EU institutions have produced translated text in it for two decades. Institutional status creates data; data creates AI capability.
Europe developed a second, more contextual instrument. The European Language Equality (ELE) project defined the Digital Language Equality (DLE) metric, which combines technological factors (available resources and tools) with situational factors of a social and economic nature. Gaspari and colleagues applied it to 89 European languages. The difference matters: Joshi counts what exists, while DLE also asks whether the surrounding ecosystem can sustain it.
The debate continues in 2026. At AmericasNLP 2026, Coleman, Coleman, and Krishnamachari proposed a Resource Abundance Notation (RAN) that reports speaker abundance, monolingual text, and parallel corpora separately instead of compressing a language into a single label. For a buyer of AI data, that separation is exactly what a statement of work should contain.
04 · Consequences
Why do AI systems fail in low-resource languages?
Large language models learn from what the web contains, and the web is linguistically lopsided. Blasi, Anastasopoulos, and Neubig (ACL 2022) documented systematic performance inequalities across the world's languages in machine translation, question answering, reading comprehension, speech synthesis, parsing, and morphological inflection, and showed that the gaps track economic, demographic, and academic factors rather than linguistic difficulty. Multilingual models extend coverage, but sharing one set of parameters across hundreds of languages does not guarantee equivalent results for each of them. Meta's No Language Left Behind project (NLLB-200, published in Nature in 2024) scaled neural machine translation to 200 languages and around 40,000 translation directions precisely by combining transfer learning with large-scale data mining, and its own evaluation shows how uneven quality remains across that range.
The subtler failure is the gap between nominal coverage and effective utility. A language can appear on a model's list of supported languages and still perform far below its neighbors, or have been evaluated only on a narrow register such as Bible text or Wikipedia. Without reliable, human-built benchmarks you cannot even tell. That is why evaluation data is not a luxury for low-resource languages: it is the only instrument that separates a claim from a capability.
Scene
You write to a regional health service chatbot in Valencian. It answers in fluent, confident central Catalan, gets two local place names wrong, and quietly switches to Spanish when your question mentions a prescription. The model "supports Catalan." Nobody measured whether it supports you.
05 · Europe
Europe's digital language equality agenda
The European Union has 24 official languages and more than 60 regional and minority languages, and it has treated technological inequality between them as a policy problem for over a decade. The sequence is worth knowing, because it explains why European public funding now flows into language data.
2012
META-NET White Papers
The META-NET white paper series assessed 30 European languages. At least 21 showed weak or no technological support in at least one of the 4 areas studied and were described as at risk of "digital extinction."
2018
European Parliament resolution
On September 11, 2018 the Parliament adopted its resolution on language equality in the digital age (2018/2028(INI)), calling for coordinated policy, funding, and infrastructure.
2021 to 2022
European Language Equality project
A consortium of 52 partners assessed the technological support of European languages, defined the DLE metric, and produced language-specific reports plus a roadmap toward digital language equality by 2030.
2023
Strategic agenda published
The results appeared as an open-access Springer volume edited by Georg Rehm and Andy Way. English remained far ahead; French, German, and Spanish formed a second group with moderate support; most other languages showed fragmentary, weak, or no support.
I co-authored the "Deep Dive: Machine Translation" chapter of European Language Equality: A Strategic Agenda for Digital Language Equality, and the conclusion I would underline from that work is practical rather than rhetorical: machine translation between well-resourced European languages had become a commodity, while quality for smaller languages, specialized domains, and non-English pairs still depended on who had invested in data. The same logic drove NTEU, the European project in which our consortium built 552 neural machine translation engines covering the 24 official EU languages, including directions for Irish, Maltese, and other languages that the Joshi taxonomy places in the lower classes.
06 · Spain as a laboratory
Is Catalan a low-resource language? Catalan, Valencian, Basque, and Galician
Spain is one of the best places in the world to see how relative the label is, because four languages with very different histories coexist inside one state and one research system. None of them fits a single category, and each has become a showcase of a different strategy.
Catalan: well organized, still fragmentary
Catalan has a strong digital community (Softcatalà, open corpora, spell checkers, free software) that researchers describe as a case of digital resilience. Yet ELE's 2022 report on the Catalan language, authored by Maite Melero, Blanca C. Figueras, Mar Rodríguez, and Marta Villegas, placed its overall technological support in the "fragmentary" category, with large differences between areas such as text, speech, and machine translation. The Generalitat de Catalunya responded with Projecte AINA, developed by the Barcelona Supercomputing Center (BSC), which has released open corpora and models for Catalan. In April 2025 the Spanish government released the ALIA family through BSC, a public, open infrastructure for Spanish and the co-official languages; its Salamandra models were trained on 35 European languages under an Apache 2.0 license.
For us at Pangeanic this is not an abstract case study. We have worked with BSC on customised datasets of bilingual segments classified by domain and style (with particular emphasis on Catalan), on data annotation with quality control in our PECAT platform, on reinforcement learning from human feedback with reward model data, and on multilingual hate speech detection datasets to reduce bias in language models. The full scope is described in our Barcelona Supercomputing Center use case.
Valencian: the variety inside the language
Resource scarcity is not uniform within a language. The same ELE report recognizes that in the Valencian Community the traditional name of the language is valencià, and it inventories Valencian-specific resources such as the Corpus Informatitzat del Valencià and the Corpus Toponímic Valencià, while warning that several corpora are small, restricted by license, or lack comparable evaluation sets. AINA's CATalog dataset states openly that most of its material is central Catalan and that Valencian and Balearic text is being actively added to reduce variety bias. In speech, the VIVES project published a corpus of more than 270 hours from the Corts Valencianes, conceived partly to compensate for the scarcity of Valencian speech data. A system trained mostly on central Catalan can formally support Catalan and still offer weaker coverage of Valencian or Balearic forms.
This is also where speech data meets public service. We have provided AI transcription for the Valencian regional parliament since 2021, and in 2024 we were awarded the contract to introduce AI transcription in the Spanish Congress, which covers interventions in Catalan, Galician, and Basque alongside Spanish. Parliamentary speech is exactly the kind of long-form, multi-speaker, code-switched audio that exposes resource gaps between varieties.
Basque: an isolate that built its own model
Basque has no close linguistic relatives, so it cannot lean on cross-lingual transfer the way Galician leans on Portuguese or Valencian on the rest of Catalan. The HiTZ Center (University of the Basque Country) answered with Latxa, a family of open models from 7 to 70 billion parameters pretrained on a new Basque corpus of 4.3 million documents and 4.2 billion tokens, released with 4 new evaluation benchmarks built from proficiency exams, reading comprehension, trivia, and public examinations. The paper received the Best Resource Paper award at ACL 2024. The lesson for everyone else: the evaluation suite was as important a contribution as the model.
Galician: transfer from a large neighbor
Galician illustrates the opposite strategy. Proxecto Nós, led by CiTIUS at the University of Santiago de Compostela with the Xunta de Galicia, has produced Galician corpora and open generative models (Carballo among them), including a Galician-Portuguese bilingual model that exploits the proximity of the two languages. Proximity helps, but it has a known side effect: models drift toward the larger neighbor's spelling and vocabulary unless native-speaker data and evaluation pull them back.
Reference table
Low-resource languages at a glance: examples, gaps, and initiatives
A comparison of representative languages across Europe, the Americas, Africa, and Asia. The status column summarizes published assessments; it is qualitative by design, because resource level depends on task and changes over time.
| Language | Where | Resource status (published assessment) | Main gaps for AI | Reference initiatives |
|---|---|---|---|---|
| Catalan | Catalonia, Valencian Community, Balearic Islands, Andorra | "Fragmentary" overall support (ELE, 2022); strong community and public investment since | Variety coverage, domain-specific data, evaluation sets, licensing | Projecte AINA, CATalog, ALIA, Softcatalà, Apertium |
| Valencian | Valencian Community | Covered formally under Catalan; fewer variety-specific corpora (ELE, 2022) | Speech data, variety-aware evaluation, local terminology | VIVES Corts Valencianes corpus (270+ hours), Corpus Informatitzat del Valencià |
| Basque | Basque Country, Navarre, French Basque Country | Language isolate; dedicated LLM and benchmarks since 2024 | Limited transfer from related languages; pretraining volume | Latxa and its evaluation suite (HiTZ), ALIA |
| Galician | Galicia | Growing open resources; benefits from proximity to Portuguese | Drift toward Portuguese or Spanish norms; native evaluation | Proxecto Nós, Carballo, ALIA |
| Irish, Maltese | Ireland; Malta | Class 2 "Hopefuls" (Joshi et al., 2020); EU official languages | Parallel data volume outside institutional domains | EU institutional translation data, NTEU engines |
| Quechua | Andean countries | Low-resource despite millions of speakers; many varieties | Transcribed speech, parallel corpora, orthographic variation | AmericasNLP shared tasks; Puno Quechua ASR resources (2026) |
| Nahuatl, Wixarika, Rarámuri, Hñähñu | Mexico | Included in the first AmericasNLP MT shared task (2021) | Parallel data, standardized orthography, segmentation tools | AmericasNLP |
| Zulu, Yoruba, Hausa | Southern and West Africa | Zulu in class 2 (Joshi et al.); fast-growing African NLP research | Annotated data, code-switching, speech in real acoustic conditions | Masakhane, MasakhaNER, African Languages Lab |
| Medumba | Cameroon | Documented resource-gathering barriers (AfricaNLP 2025) | Writing and encoding methods, technical terminology, funding | Community and academic documentation projects |
| Bhojpuri | India, Nepal | Class 1 "Scraping-Bys" despite tens of millions of speakers | Almost no labeled data; confusion with Hindi | Indian language data initiatives |
| Lao | Laos | Class 2 "Hopefuls" (Joshi et al.) | Script written without spaces between words: tokenization and segmentation | Southeast Asian data collection projects |
| Cebuano | Philippines | Class 3 "Rising Stars"; very large but largely bot-generated Wikipedia | Natural, human-written text beyond templated articles | Wikipedia quality audits (Tatariya et al., 2024) |
Sources: Joshi et al. (ACL 2020); ELE Report on the Catalan Language (2022); Etxaniz et al. (ACL 2024); Mager et al. (AmericasNLP 2021); Moteu Ngoli et al. (AfricaNLP 2025); Huaman et al. (AmericasNLP 2026); Tatariya et al. (2024). Full references at the end of the article.
07 · Beyond Europe
Africa, the Americas, and Asia
Africa: where most of the diversity is
Africa concentrates a very large share of the world's linguistic diversity, and it remains thinly represented in corpora and models. A review of 884 African NLP papers published between 2020 and 2025 (Alabi, Hedderich, Adelani, and Klakow, EMNLP 2025) describes sustained growth in research, with very uneven coverage across languages, tasks, and regions. Masakhane, founded in 2019 as a distributed research community, changed the method as much as the output: its participatory machine translation study (Nekoto et al., 2020) produced data and systems for more than 30 African languages with local researchers and speakers in charge. The obstacles are not only volume. For Medumba, spoken in Cameroon, researchers documented problems with writing and encoding methods, the absence of standardized technical terminology, scarce digital resources, and limited funding. When we collect African datasets for AI training, code-switching and real acoustic conditions (markets, transport, rural settings) turn out to matter as much as vocabulary.
The Americas: Indigenous languages on their own terms
In Latin America, the technological position of Spanish and Portuguese contrasts sharply with that of Indigenous languages. The difficulty rarely comes from a small number of speakers: it comes from missing digital corpora, orthographies that are not always stable, and scarce segmentation tools, parallel data, recorded speech, and evaluation sets (Mager et al., COLING 2018). The AmericasNLP workshop, started in 2021 within the Association for Computational Linguistics, organized the first shared task on open machine translation for 10 Indigenous languages paired with Spanish: Aymara, Bribri, Asháninka, Guaraní, Wixarika, Nahuatl, Hñähñu, Quechua, Shipibo-Konibo, and Rarámuri. Later editions expanded to more languages, to educational materials, and to translation metrics. In 2026 a community-centered project built speech recognition resources for Puno Quechua from dozens of hours of recordings and transcriptions, with the community involved in design and collection. That model of work (communities setting goals and validating outputs) is becoming the standard, not the exception.
Asia: scripts, segmentation, and scale
Asia shows that low-resource does not mean small. Bhojpuri appears in Joshi's class 1 despite a speaker base of tens of millions; Konkani and Lao sit in class 2; Cebuano and Indonesian in class 3. Here the gaps are often technical before they are statistical: scripts written without spaces between words (Lao, Thai, Khmer) break naive tokenizers, and closely related languages are frequently mislabeled as the dominant one, which inflates apparent coverage for the larger language and erases the smaller one. My reading, and it is an inference rather than a published measurement, is that the next wave of useful Asian language AI will come from carefully segmented, correctly labeled speech and text collected in-country, not from larger crawls. That is the principle behind our Southeast Asian datasets for AI training.
08 · Techniques
How do you build AI for a low-resource language?
The literature (Hedderich et al., 2021; Haddow et al., 2022; Ranathunga et al., 2023) converges on 6 families of techniques. None is universal, and each carries a specific risk.
01
Transfer learning
Fine-tune a multilingual pretrained model (mBERT, XLM-R, mT5, BLOOM, or open LLMs) on the target language.
Risk: weak when the target is distant from the languages that dominate pretraining.
02
Multilingual MT and related languages
Train jointly with linguistically related, better-resourced languages so that the smaller one benefits from shared structure.
Risk: drift toward the larger neighbor's norms.
03
Data augmentation
Synthetic data and back-translation create extra training pairs from monolingual text.
Risk: synthetic data amplifies the errors of the system that produced it.
04
Semi-supervised and distant supervision
Exploit unlabeled text, rules, and dictionaries to generate weak labels at scale.
Risk: noisy labels require human-validated test sets to detect.
05
Linguistic knowledge and rules
Lexicons, grammars, and morphological analyzers when data cannot support a purely statistical approach. Apertium, the open-source rule-based MT platform started in 2005 at the Universitat d'Alacant, remains a reference for related-language pairs.
Risk: expensive expert time; limited coverage of open domains.
06
Participatory data collection
Speakers contribute, validate, and govern text and voice data, as in Mozilla Common Voice, Masakhane, or AmericasNLP projects.
Risk: needs time, coordination, fair compensation, and real consent mechanisms.
09 · The critique
Is "low-resource" the right label?
The NLP community still uses the term as standard, but its critics have sharpened what it hides. Steven Bird argues that the label lumps together very different situations (standardized languages without technological support, local languages of mainly oral tradition, contact languages) whose social functions and technological expectations are not comparable, and he asks whether NLP must be extractive at all (ACL 2024). Treating linguistic data as raw material to be mined can reproduce colonial relationships and sideline community authorities. Mager, Mager, Kann, and Vu (ACL 2023) interviewed community leaders, teachers, and activists and concluded that including speakers is essential for useful and ethical machine translation of Indigenous languages, which points toward linguistic data sovereignty: communities keep control over how the materials they contribute are used to train AI.
I agree with the critique, and I think it has a commercial translation that the industry has been slow to accept. Data whose provenance, consent, and license cannot be documented is a liability, and in low-resource languages that liability is concentrated, because the same few sources get reused everywhere. Ethical sourcing is not a tax on low-resource AI; it is the only way to build datasets that survive audit. This is why we insist on a human-in-the-center model, where native speakers are not a validation afterthought but the people who decide what "correct" means for their variety.
10 · Decision criteria
7 questions to ask before training or buying AI for a low-resource language
Whether you are fine-tuning a sovereign model, commissioning a dataset, or procuring machine translation for a public service, these questions separate nominal coverage from real capability.
1. Which variety, register, and domain?
"Catalan" or "Quechua" is not a specification. Name the variety, the audience, and the domain, and require data from each.
2. What does "supported" mean?
Ask for the benchmark, the test set, the metric, and the date. A language on a list is a claim; a score on a human-validated test set is evidence.
3. Where did the data come from?
Demand documented provenance, license, and consent for every source, especially community-contributed speech and text.
4. How much is web-crawled, and who audited it?
Crawled corpora for smaller languages contain mislabeled and unusable text. Native-speaker audits catch what automatic filters miss.
5. Is there a held-out evaluation set built by humans?
Without one you cannot compare vendors or detect regressions between model versions.
6. Is synthetic data labeled and separated?
Back-translated and generated data must be tagged so that it never contaminates evaluation sets.
7. Who validates outputs in production?
Define the native-speaker review loop, the quality thresholds, and how feedback returns to training.
When references do not exist
Reference-free machine translation quality estimation (MTQE) predicts quality without a human reference, which makes it especially useful where reference translations are scarce.
11 · Our work
How we approach low-resource languages at Pangeanic
Pangeanic began by building and processing bilingual data for machine translation systems, long before "training data" became a market category. That history taught us that the hard part of a low-resource language is never the model architecture: it is finding, licensing, cleaning, aligning, and validating the data, and then proving with evaluation that the result works for the people who speak it. Today that work runs through our Data for AI services (sourcing, collection, annotation, RLHF, and evaluation data), our parallel corpora [CONFIRMAR URL canónica], and our AI evaluation practice, with native-speaker validation at each step. The same data discipline sits behind our Deep Adaptive AI Translation, which adapts output to a client's terminology and style, a need that is sharper, not softer, in languages where generic engines have seen little text.
If you want a longer view of where machine translation still struggles, I wrote about the languages that defy machine translation from our production experience.
Documented and inferred
The figures in this article (class sizes, audit results, corpus sizes, hours of speech, project dates) come from the published sources listed below. The ELE assessment of Catalan dates from 2022 and predates the release of ALIA in 2025, so Catalan's current position is likely stronger than that snapshot; no comparable reassessment has been published yet. The gap between listed and effective support for Valencian, the drift of Galician models toward Portuguese, and the prediction about Asian data collection are my own assessments from production work, not measured results. Joshi's taxonomy reflects the catalogs available in 2020 and should be read as a snapshot, not a ranking.
FAQ
Frequently asked questions about low-resource languages
What is a low-resource language?
A low-resource language is a language with too little machine-readable text, speech, parallel data, annotated data, tooling, or evaluation benchmarks to build language technology that performs as well as it does for English. The label describes a data condition that depends on the task, the domain, and the moment, not the richness of the language.
Is a low-resource language the same as a minority or endangered language?
No. A minority language is defined by its sociopolitical position, an endangered language by weakened intergenerational transmission, and a low-resource language by the scarcity of digital data and tools. The categories often overlap, but a minority language can be well resourced and a language with tens of millions of speakers can be low-resource.
How many languages are low-resource?
Most of them. In the taxonomy of Joshi et al. (ACL 2020), 2,191 languages fell into the lowest class, with virtually no labeled or unlabeled data, while only 7 languages (including English, Spanish, German, Japanese, and French) formed the top class. More than 7,000 languages are spoken worldwide.
Is Catalan a low-resource language?
It depends on the task. The European Language Equality report of 2022 rated Catalan's overall technological support as "fragmentary," with large differences between text, speech, and machine translation. Public programs such as Projecte AINA and ALIA have since added open corpora and models, but variety coverage (Valencian, Balearic) and evaluation data remain gaps.
Are Basque and Galician low-resource languages?
Both have historically been under-resourced for AI and have followed different strategies. Basque, a language isolate, now has Latxa, a family of open models with a dedicated evaluation suite from the HiTZ Center. Galician benefits from transfer from Portuguese through Proxecto Nós and its Carballo models, though models can drift toward Portuguese norms without native-speaker evaluation.
Why does machine translation perform worse in low-resource languages?
Neural machine translation and large language models learn from parallel and monolingual data, and low-resource languages have little of it, often concentrated in narrow domains and mixed with mislabeled or low-quality text. Without human-built test sets, errors also go undetected, so a language can appear as supported while performing far below better-resourced languages.
How do you build AI for a low-resource language?
The main techniques are transfer learning from multilingual models, joint training with related languages, data augmentation such as back-translation, semi-supervised learning, rule-based and lexical resources, and participatory data collection with speaker communities. All of them depend on high-quality, well-documented data and on human-validated evaluation sets.
What data does Pangeanic provide for low-resource languages?
Pangeanic sources, collects, annotates, and validates text, speech, and parallel data for AI training and evaluation, including RLHF and preference data, with native-speaker validation. Its work includes datasets for the Barcelona Supercomputing Center with emphasis on Catalan, AI transcription covering Catalan, Galician, and Basque in the Spanish Congress, and data collection for African and Southeast Asian languages.
Data for AI
Building AI for a language the web forgot?
Tell us the language, variety, and domain. We will tell you what data exists, what has to be collected, and how to prove the model works.
References
Sources
- Alabi, J. O., Hedderich, M. A., Adelani, D. I., and Klakow, D. (2025). Charting the Landscape of African NLP: Mapping Progress and Shaping the Road Ahead. Proceedings of EMNLP 2025, 27807-27841. doi:10.18653/v1/2025.emnlp-main.1414
- Bird, S. (2024). Must NLP be Extractive? Proceedings of ACL 2024, 14915-14929. doi:10.18653/v1/2024.acl-long.797
- Blasi, D., Anastasopoulos, A., and Neubig, G. (2022). Systematic Inequalities in Language Technology Performance across the World's Languages. Proceedings of ACL 2022, 5486-5505. doi:10.18653/v1/2022.acl-long.376
- Coleman, J., Coleman, T., and Krishnamachari, B. (2026). RAN: Resource Abundance Notation for Languages in NLP. Proceedings of AmericasNLP 2026, 168-172.
- Etxaniz, J., Sainz, O., Perez, N., et al. (2024). Latxa: An Open Language Model and Evaluation Suite for Basque. Proceedings of ACL 2024.
- Gaspari, F., et al. (2023). Digital Language Equality: Definition, Metric, Dashboard. In Rehm and Way (eds.), European Language Equality, Springer, 39-73.
- Haddow, B., Bawden, R., Miceli Barone, A. V., Helcl, J., and Birch, A. (2022). Survey of Low-Resource Machine Translation. Computational Linguistics 48(3), 673-732. doi:10.1162/coli_a_00446
- Hedderich, M. A., Lange, L., Adel, H., Strötgen, J., and Klakow, D. (2021). A Survey on Recent Approaches for Natural Language Processing in Low-Resource Scenarios. Proceedings of NAACL 2021, 2545-2568. doi:10.18653/v1/2021.naacl-main.201
- Huaman, E., Gamarra Lafuente, A., Cordova, J., and Korhonen, A. (2026). Building Community-Centred NLP Resources for Puno Quechua. Proceedings of AmericasNLP 2026.
- Joshi, P., Santy, S., Budhiraja, A., Bali, K., and Choudhury, M. (2020). The State and Fate of Linguistic Diversity and Inclusion in the NLP World. Proceedings of ACL 2020, 6282-6293. doi:10.18653/v1/2020.acl-main.560
- Kreutzer, J., Caswell, I., et al. (2022). Quality at a Glance: An Audit of Web-Crawled Multilingual Datasets. Transactions of the ACL 10, 50-72. doi:10.1162/tacl_a_00447
- Mager, M., Gutierrez-Vasques, X., Sierra, G., and Meza-Ruiz, I. (2018). Challenges of Language Technologies for the Indigenous Languages of the Americas. Proceedings of COLING 2018, 55-69.
- Mager, M., et al. (2021). Findings of the AmericasNLP 2021 Shared Task on Open Machine Translation for Indigenous Languages of the Americas. Proceedings of AmericasNLP 2021, 202-217.
- Mager, M., Mager, E., Kann, K., and Vu, N. T. (2023). Ethical Considerations for Machine Translation of Indigenous Languages: Giving a Voice to the Speakers. Proceedings of ACL 2023, 4871-4897.
- Melero, M., C. Figueras, B., Rodríguez, M., and Villegas, M. (2022). Report on the Catalan Language. European Language Equality, Deliverable D1.6.
- Moteu Ngoli, T., Christabel, M., and Yopa, N. (2025). Challenges and Limitations in Gathering Resources for Low-Resource Languages: The Case of Medumba. Proceedings of AfricaNLP 2025, 136-142.
- Nekoto, W., et al. (2020). Participatory Research for Low-resourced Machine Translation: A Case Study in African Languages. Findings of EMNLP 2020, 2144-2160.
- NLLB Team (2024). Scaling neural machine translation to 200 languages. Nature 630, 841-846. doi:10.1038/s41586-024-07335-x
- Ranathunga, S., et al. (2023). Neural Machine Translation for Low-resource Languages: A Survey. ACM Computing Surveys 55(11). doi:10.1145/3567592
- Rehm, G., and Way, A. (eds.) (2023). European Language Equality: A Strategic Agenda for Digital Language Equality. Springer, Cognitive Technologies. doi:10.1007/978-3-031-28819-7
- Tatariya, K., et al. (2024). How Good is Your Wikipedia? Auditing Data Quality for Low-resource and Multilingual NLP. arXiv:2411.05527

