608 languages in 37 MB: building African-first language identification

Language identification sounds solved. It is the first step of almost every multilingual pipeline: before you
can translate, filter a web crawl, route a support ticket or pick a speech model, you need to know what language
the text is in. Open models like GlotLID and Meta’s fastText LID cover hundreds of languages, and on long,
clean paragraphs they are very good.
Real input is not long and clean. It is “Ẹ kú àárọ̀”, “how far”, a product review with an emoji, a WhatsApp
message with no tone marks, a URL. And for African languages, the gaps are wide: Kinyarwanda, Xhosa, Wolof and
Lingala are routinely confused with their neighbours, and Nigerian Pidgin is often read as English.
So we built two models for 608 languages, African-first, designed for the input people actually type:
•lid-lite-608: a 37 MB fastText model, about 6,800 texts per second on one CPU thread.
•lid-neural-608: a 140M-parameter ModernBERT classifier fine-tuned from mmBERT-small, the most accurate model we tested on short and conversational text.
Both are Apache-2.0 and trained only on data that allows commercial use.
The results
We scored every model on the same data, mapping each model’s language codes to ours so that naming
differences never count as errors.

| lid-neural-608 | lid-lite-608 | GlotLID v3 | Meta fastText LID-218 | |
|---|---|---|---|---|
| Size | 140M params | 37 MB | 1,687 MB | 1,176 MB |
| Held-out test, accuracy (240k samples, 608 languages) | 0.867 | 0.858 | 0.783 | 0.280 |
| Short input (1-5 words), web-weighted, traffic mode | 0.906 | 0.829 | 0.808 | 0.764 |
| Everyday phrases (“good morning”, “sign in”, …) | 0.943 | 0.838 | 0.857 | 0.867 |
| Tatoeba (conversational, 34k) | 0.930 | 0.922 | 0.868 | 0.557 |
| UDHR-LID (291 languages) | 0.914 | 0.910 | 0.908 | 0.567 |
| FLORES+ devtest (202 languages) | 0.942 | 0.934 | 0.957 | 0.874 |
On the 20 major African languages we track, lid-lite-608 averages 0.839 F1 against GlotLID’s 0.757, with
the biggest gains where it matters most: Kinyarwanda (0.758 vs 0.371), Xhosa (0.840 vs 0.602), Wolof (0.809 vs
0.600) and Lingala (0.814 vs 0.642).

FLORES+ is the one benchmark where GlotLID leads. Its sentences are long, carefully translated news and wiki
text, the kind of input every model handles well. The gap opens on short and conversational text: on the
first 1-5 words of the same FLORES+ sentences, both of our models are ahead.
One model, two ways to read it
The most useful thing we added is not a bigger model but a second way of reading the one we have.
A language identifier trained on balanced data treats all 608 languages as equally likely. That is exactly
right when you are mining a web crawl for low-resource text: you want every Lingala sentence, and you want it
labelled Lingala even when it is short. But it is wrong for a chat box. If a user types “Bonjour mon ami”,
the overwhelmingly likely answer is French, not one of the dozens of small languages that share the same
words.
So both models ship with two modes:
•coverage (default): every language equally likely. Best per-language accuracy, for corpus building.
•traffic: each language’s score is shifted by how common it is in real text. Formally, we add 0.2 × log(real-world share / training share) to each language’s log-probability, using frequencies we measured on the web. Short and ambiguous input then leans towards the languages that actually dominate traffic.
from olaverse import LIDLite608
LIDLite608().predict("Bonjour mon ami") # 'dhv_Latn' (coverage)
LIDLite608(mode="traffic").predict("Bonjour mon ami") # 'fra_Latn' (traffic)
Traffic mode adds 6.6 points of accuracy on 1-5 word input for lid-lite-608 and 5.5 for lid-neural-608, and it
costs nothing: the same weights, one extra vector.
The bug that taught us the most: “Good morning” was not English
An early version of lid-lite scored well on every benchmark we had, and then we typed “Good morning” into it.
It did not come back as English, and neither did most other everyday English phrases: precision for English on
short phrases was 0.19.
The cause was in the data, not the model. Web text in small languages is full of English: menus, cookie
banners, “Read more”, “Sign in”, quoted headlines. Each of those English fragments was labelled with the
language of the page it came from, so the model learned that short English phrases are evidence for hundreds
of other languages.
The fix had four parts:
1.A foreign-text filter that removes boilerplate and lines in a different language from their page label before training.
2.Rebalancing the major languages, so English, French, Spanish and the other high-traffic languages were not drowned out by the long tail.
3.Tatoeba sentences, about 1.6M samples of short, conversational, human-written text in hundreds of languages.
4.A probe test of 107 everyday phrases in 11 common languages (greetings, “sign in”, “add to cart”), removed from the training data and checked on every build.
The probe test now sits next to the benchmarks in every evaluation: lid-neural-608 gets 0.943 of the phrases
right in traffic mode, the best of all the models we tested. The lesson generalises: aggregate benchmarks
can hide a failure that every user hits in their first minute. Build a small test of the obvious cases and
never ship without it.
Built for messy input
Both models were trained on input that looks like real input:
•Typos, missing diacritics, case changes and emoji were added to a share of the training samples, so “se daadaa ni” is recognised as Yorùbá as reliably as “ṣé dáadáa ni”.
•Four length buckets per language, from 1-5 words to full paragraphs, because a model trained only on paragraphs fails on queries.
•A noise class (zxx_Zxxx) for numbers, URLs, code and other non-language text, so a stray order number is not reported as a language. lid-lite-608 reaches 0.961 F1 on it, against 0.620 for GlotLID.
•Data hygiene: we removed religious text (which dominates many low-resource corpora and skews the vocabulary), text in the wrong script for its label, and mislabelled documents.
The held-out test was built from websites never seen in training, 100 samples per length bucket per
language, so the scores reflect new sources rather than memorised ones.
Which one to use
| lid-lite-608 | lid-neural-608 | |
|---|---|---|
| Size | 37 MB | 140M parameters (~560 MB) |
| Speed | ~6,800 texts/s on one CPU thread | ~1,100 texts/s on an A100 |
| Best for | bulk filtering, CPU and edge deployment | user input, short text, highest accuracy |
Or use both: run lid-lite-608 as a fast first pass over everything, and send only short or low-confidence
inputs to lid-neural-608.
Try it
pip install "olaverse[lid]" # lid-lite-608
pip install "olaverse[deeplearning]" # lid-neural-608
from olaverse import LIDLite608, LIDNeural608
LIDLite608().predict("Ẹ kú àárọ̀, ṣé dáadáa ni?") # 'yor_Latn'
lid = LIDNeural608(mode="traffic")
lid.predict_batch(["Good morning", "Mo fẹ́ lọ sí ọjà", "Habari za asubuhi"])
# ['eng_Latn', 'yor_Latn', 'swh_Latn']
Both models also work with plain fasttext and transformers; the model cards have the code. Labels are ISO
639-3 plus script (yor_Latn, srp_Cyrl), so the same language in two scripts is never confused.
•Models: olaverse/lid-lite-608, olaverse/lid-neural-608
•Library: olaverse on PyPI · docs
If you work with a language we get wrong, we would like to hear about it: open a discussion on either model
page with a few example sentences.
Enjoyed this? Every model we write about is open weight and free to download.


