Find the low-resource language text hiding in a web crawl

    Sort billions of web documents into 608 languages on CPUs, drop the junk, and pull out the African and low-resource text that general detectors misfile.

    ML & data teamsResearch labsLanguage archivesLocalisation
    0.839
    AVG F1, 20 AFRICAN LANGUAGES (GLOTLID 0.757)
    37 MB
    MODEL SIZE
    0.961
    F1 ON THE NOISE CLASS
    / The problem

    Why this breaks today

    Training data for a low-resource language mostly exists, scattered through web crawls, but it is mislabelled. Detectors with a small label set force Kinyarwanda, Wolof or Lingala into whichever bigger language looks closest, and the text is lost.

    Crawls are also full of text that is no language at all: numbers, URLs, code. Pass it through a detector without a junk class and it gets scattered across real languages as noise.

    / How it works

    The pipeline, step by step

    01

    Run the lite model in coverage mode

    Coverage mode treats every language as equally likely, which is what you want when the point is finding rare ones. lid-lite-608 does it in 37 MB at about 6,800 texts per second per CPU thread: 3× the languages of Meta’s fastText LID at 1/32 of its size.

    lid-lite-608
    02

    Split long pages and aggregate

    Classify ~256-word chunks and take the majority. This also catches pages that switch language partway through, which a single whole-page call would average away.

    lid-lite-608
    03

    Drop the junk class

    zxx_Zxxx marks numbers, URLs, code and other non-language text, at F1 0.961 for the lite model. Filter it out before anything else sees it.

    lid-lite-608
    04

    Second opinion on the hard cases

    Send short or low-confidence chunks to lid-neural-608, the most accurate model in our tests (0.867 accuracy, macro-F1 0.862 on the 240k-sample held-out set), for the final label.

    lid-neural-608
    / The code

    Running in about ten lines

    The Olaverse SDK wraps the models with sane defaults. If you would rather not add a dependency, the second tab is the same pipeline in plain transformers.

    $pip install "olaverse[lid]"
    mine.py
    from collections import Counter
    from olaverse import LIDLite608
    
    lid = LIDLite608()            # mode="coverage" (default): every language equally likely
    
    def chunks(text, n=256):
        words = text.split()
        return [" ".join(words[i:i + n]) for i in range(0, len(words), n)] or [text]
    
    def label_page(text):
        labels = [l for l in lid.predict_batch(chunks(text)) if l != "zxx_Zxxx"]
        return Counter(labels).most_common(1)[0][0] if labels else None
    
    buckets = {}
    for page in crawl:
        lang = label_page(page)
        if lang:
            buckets.setdefault(lang, []).append(page)
    
    buckets.get("kin_Latn", [])[:3]    # the Kinyarwanda pages

    Language coverage

    / 20 AFRICAN LANGUAGES BENCHMARKED

    The models cover 608 languages across 36 scripts. These are the 20 African languages the model card benchmarks one by one against GlotLID.

    Yorùbá
    yo · yor
    Hausa
    ha · hau
    Igbo
    ig · ibo
    Nigerian Pidgin
    — · pcm
    Swahili
    sw · swh
    Amharic
    am · amh
    Zulu
    zu · zul
    Xhosa
    xh · xho
    Oromo
    om · gaz
    Luganda
    lg · lug
    Kinyarwanda
    rw · kin
    Kirundi
    rn · run
    Lingala
    ln · lin
    Twi
    tw · twi
    Ewe
    ee · ewe
    Fon
    — · fon
    Wolof
    wo · wol
    Somali
    so · som
    Shona
    sn · sna
    Tigrinya
    ti · tir

    Average F1 0.839 across these 20 against 0.757 for GlotLID, with the biggest gains on Kinyarwanda (0.758 vs 0.371), Xhosa, Wolof and Lingala.

    / This is you if
    • You are building a training corpus for a language that general crawls under-represent.
    • You filter Common Crawl-scale data and need it to run on CPUs.
    • Your current detector has a small label set and dumps rare languages into big ones.

    Every figure on this page comes from the published model cards. Download the weights and check us.

    Want this on your data?

    The weights are open, so you can build it yourself this afternoon. If you would rather we tuned it to your domain, evaluated it properly, and handed it over, tell us what you are working on.