Models Oct 06, 2026 5 min read

    608 languages in 37 MB: building African-first language identification

    OO
    Olumide Ola
    Olaverse Lab
    608 languages in 37 MB: building African-first language identification

    Language identification sounds solved. It is the first step of almost every multilingual pipeline: before you
    can translate, filter a web crawl, route a support ticket or pick a speech model, you need to know what language
    the text is in. Open models like GlotLID and Meta’s fastText LID cover hundreds of languages, and on long,
    clean paragraphs they are very good.

    Real input is not long and clean. It is “Ẹ kú àárọ̀”, “how far”, a product review with an emoji, a WhatsApp
    message with no tone marks, a URL. And for African languages, the gaps are wide: Kinyarwanda, Xhosa, Wolof and
    Lingala are routinely confused with their neighbours, and Nigerian Pidgin is often read as English.

    So we built two models for 608 languages, African-first, designed for the input people actually type:

    •lid-lite-608: a 37 MB fastText model, about 6,800 texts per second on one CPU thread.

    •lid-neural-608: a 140M-parameter ModernBERT classifier fine-tuned from mmBERT-small, the most accurate model we tested on short and conversational text.

    Both are Apache-2.0 and trained only on data that allows commercial use.

    The results

    We scored every model on the same data, mapping each model’s language codes to ours so that naming
    differences never count as errors.

    lid-neural-608 vs lid-lite-608, GlotLID and Meta
    lid-neural-608lid-lite-608GlotLID v3Meta fastText LID-218
    Size140M params37 MB1,687 MB1,176 MB
    Held-out test, accuracy (240k samples, 608 languages)0.8670.8580.7830.280
    Short input (1-5 words), web-weighted, traffic mode0.9060.8290.8080.764
    Everyday phrases (“good morning”, “sign in”, …)0.9430.8380.8570.867
    Tatoeba (conversational, 34k)0.9300.9220.8680.557
    UDHR-LID (291 languages)0.9140.9100.9080.567
    FLORES+ devtest (202 languages)0.9420.9340.9570.874

    On the 20 major African languages we track, lid-lite-608 averages 0.839 F1 against GlotLID’s 0.757, with
    the biggest gains where it matters most: Kinyarwanda (0.758 vs 0.371), Xhosa (0.840 vs 0.602), Wolof (0.809 vs
    0.600) and Lingala (0.814 vs 0.642).

    African languages F1

    FLORES+ is the one benchmark where GlotLID leads. Its sentences are long, carefully translated news and wiki
    text, the kind of input every model handles well. The gap opens on short and conversational text: on the
    first 1-5 words of the same FLORES+ sentences, both of our models are ahead.

    One model, two ways to read it

    The most useful thing we added is not a bigger model but a second way of reading the one we have.

    A language identifier trained on balanced data treats all 608 languages as equally likely. That is exactly
    right when you are mining a web crawl for low-resource text: you want every Lingala sentence, and you want it
    labelled Lingala even when it is short. But it is wrong for a chat box. If a user types “Bonjour mon ami”,
    the overwhelmingly likely answer is French, not one of the dozens of small languages that share the same
    words.

    So both models ship with two modes:

    •coverage (default): every language equally likely. Best per-language accuracy, for corpus building.

    •traffic: each language’s score is shifted by how common it is in real text. Formally, we add 0.2 × log(real-world share / training share) to each language’s log-probability, using frequencies we measured on the web. Short and ambiguous input then leans towards the languages that actually dominate traffic.

    from olaverse import LIDLite608
    
    LIDLite608().predict("Bonjour mon ami")                  # 'dhv_Latn'  (coverage)
    LIDLite608(mode="traffic").predict("Bonjour mon ami")    # 'fra_Latn'  (traffic)

    Traffic mode adds 6.6 points of accuracy on 1-5 word input for lid-lite-608 and 5.5 for lid-neural-608, and it
    costs nothing: the same weights, one extra vector.

    The bug that taught us the most: “Good morning” was not English

    An early version of lid-lite scored well on every benchmark we had, and then we typed “Good morning” into it.
    It did not come back as English, and neither did most other everyday English phrases: precision for English on
    short phrases was 0.19.

    The cause was in the data, not the model. Web text in small languages is full of English: menus, cookie
    banners, “Read more”, “Sign in”, quoted headlines. Each of those English fragments was labelled with the
    language of the page it came from, so the model learned that short English phrases are evidence for hundreds
    of other languages.

    The fix had four parts:

    1.A foreign-text filter that removes boilerplate and lines in a different language from their page label before training.

    2.Rebalancing the major languages, so English, French, Spanish and the other high-traffic languages were not drowned out by the long tail.

    3.Tatoeba sentences, about 1.6M samples of short, conversational, human-written text in hundreds of languages.

    4.A probe test of 107 everyday phrases in 11 common languages (greetings, “sign in”, “add to cart”), removed from the training data and checked on every build.

    The probe test now sits next to the benchmarks in every evaluation: lid-neural-608 gets 0.943 of the phrases
    right in traffic mode, the best of all the models we tested. The lesson generalises: aggregate benchmarks
    can hide a failure that every user hits in their first minute.
     Build a small test of the obvious cases and
    never ship without it.

    Built for messy input

    Both models were trained on input that looks like real input:

    •Typos, missing diacritics, case changes and emoji were added to a share of the training samples, so “se daadaa ni” is recognised as Yorùbá as reliably as “ṣé dáadáa ni”.

    •Four length buckets per language, from 1-5 words to full paragraphs, because a model trained only on paragraphs fails on queries.

    •A noise class (zxx_Zxxx) for numbers, URLs, code and other non-language text, so a stray order number is not reported as a language. lid-lite-608 reaches 0.961 F1 on it, against 0.620 for GlotLID.

    •Data hygiene: we removed religious text (which dominates many low-resource corpora and skews the vocabulary), text in the wrong script for its label, and mislabelled documents.

    The held-out test was built from websites never seen in training, 100 samples per length bucket per
    language, so the scores reflect new sources rather than memorised ones.

    Which one to use

    lid-lite-608lid-neural-608
    Size37 MB140M parameters (~560 MB)
    Speed~6,800 texts/s on one CPU thread~1,100 texts/s on an A100
    Best forbulk filtering, CPU and edge deploymentuser input, short text, highest accuracy

    Or use both: run lid-lite-608 as a fast first pass over everything, and send only short or low-confidence
    inputs to lid-neural-608.

    Try it

    pip install "olaverse[lid]"            # lid-lite-608
    pip install "olaverse[deeplearning]"   # lid-neural-608
    from olaverse import LIDLite608, LIDNeural608
    
    LIDLite608().predict("Ẹ kú àárọ̀, ṣé dáadáa ni?")     # 'yor_Latn'
    
    lid = LIDNeural608(mode="traffic")
    lid.predict_batch(["Good morning", "Mo fẹ́ lọ sí ọjà", "Habari za asubuhi"])
    # ['eng_Latn', 'yor_Latn', 'swh_Latn']

    Both models also work with plain fasttext and transformers; the model cards have the code. Labels are ISO
    639-3 plus script (yor_Latn, srp_Cyrl), so the same language in two scripts is never confused.

    •Models: olaverse/lid-lite-608, olaverse/lid-neural-608

    •Library: olaverse on PyPI · docs

    If you work with a language we get wrong, we would like to hear about it: open a discussion on either model
    page with a few example sentences.

    Enjoyed this? Every model we write about is open weight and free to download.

    🤗 Follow on Hugging Face

    More from the Blog

    diacnet-2.0: how we fixed Yorùbá, added Arabic, and removed 90% of our Hausa errors with one function
    Models

    diacnet-2.0: how we fixed Yorùbá, added Arabic, and removed 90% of our Hausa errors with one function

    diactag-2.0: Arabic vowel marks from a model that cannot corrupt your text
    Models

    diactag-2.0: Arabic vowel marks from a model that cannot corrupt your text

    A diacritic model that cannot corrupt your text
    Models

    A diacritic model that cannot corrupt your text