Sort billions of web documents into 608 languages on CPUs, drop the junk, and pull out the African and low-resource text that general detectors misfile.
Training data for a low-resource language mostly exists, scattered through web crawls, but it is mislabelled. Detectors with a small label set force Kinyarwanda, Wolof or Lingala into whichever bigger language looks closest, and the text is lost.
Crawls are also full of text that is no language at all: numbers, URLs, code. Pass it through a detector without a junk class and it gets scattered across real languages as noise.
Coverage mode treats every language as equally likely, which is what you want when the point is finding rare ones. lid-lite-608 does it in 37 MB at about 6,800 texts per second per CPU thread: 3× the languages of Meta’s fastText LID at 1/32 of its size.
lid-lite-608Classify ~256-word chunks and take the majority. This also catches pages that switch language partway through, which a single whole-page call would average away.
lid-lite-608zxx_Zxxx marks numbers, URLs, code and other non-language text, at F1 0.961 for the lite model. Filter it out before anything else sees it.
lid-lite-608Send short or low-confidence chunks to lid-neural-608, the most accurate model in our tests (0.867 accuracy, macro-F1 0.862 on the 240k-sample held-out set), for the final label.
lid-neural-608The Olaverse SDK wraps the models with sane defaults. If you would rather not add a dependency, the second tab is the same pipeline in plain transformers.
from collections import Counter
from olaverse import LIDLite608
lid = LIDLite608() # mode="coverage" (default): every language equally likely
def chunks(text, n=256):
words = text.split()
return [" ".join(words[i:i + n]) for i in range(0, len(words), n)] or [text]
def label_page(text):
labels = [l for l in lid.predict_batch(chunks(text)) if l != "zxx_Zxxx"]
return Counter(labels).most_common(1)[0][0] if labels else None
buckets = {}
for page in crawl:
lang = label_page(page)
if lang:
buckets.setdefault(lang, []).append(page)
buckets.get("kin_Latn", [])[:3] # the Kinyarwanda pagesThe models cover 608 languages across 36 scripts. These are the 20 African languages the model card benchmarks one by one against GlotLID.
Average F1 0.839 across these 20 against 0.757 for GlotLID, with the biggest gains on Kinyarwanda (0.758 vs 0.371), Xhosa, Wolof and Lingala.
Every figure on this page comes from the published model cards. Download the weights and check us.
Detect the language of support tickets, chat messages, queries, and documents across 608 languages, built African-first and accurate on the short text real products see.
/ ML engineeringGenerate verified (query, passage) pairs from your own corpus in 25 languages, and fine-tune retrievers and rerankers on data that matches your domain.
The weights are open, so you can build it yourself this afternoon. If you would rather we tuned it to your domain, evaluated it properly, and handed it over, tell us what you are working on.