608 languages in 37 MB: building African-first language identification
Language identification sounds solved. It is the first step of almost every multilingual pipeline: before youcan translate, filter a web crawl, route a support ticket or pick a speech model, you need to know what languagethe text is in. Open models like GlotLID and Meta's fastText LID cover hundreds of languages, and on long,clean paragraphs [...]
diacnet-2.0: how we fixed Yorùbá, added Arabic, and removed 90% of our Hausa errors with one function
diacnet restores diacritics by rewriting text: a byte-levelByT5 model reads <yor> se eranko naa si gbo o? and writes ṣé ẹranko náà sì gbọ́ ọ?. Today we are releasingdiacnet-2.0 (582M) anddiacnet-mini-2.0 (300M): 11 languages including Arabic,Yorùbá error cut by 60%, and, with one small post-processing step, lower error than our character taggerdiactag-2.0 on 8 of 10 languages. This post is mostly about [...]
diactag-2.0: Arabic vowel marks from a model that cannot corrupt your text
Diacritics carry meaning. In Yorùbá, ọkọ́ is a hoe, ọkọ̀ a vehicle and ọkọ a husband. In Arabic, theshort vowels that distinguish kataba (he wrote) from kutiba (it was written) are usually not written atall. Text without its marks is ambiguous for readers, for text-to-speech, and for every model downstream. diactag-1.0 restored diacritics in ten languages with an unusualguarantee: it cannot change your text. It never generates [...]
A diacritic model that cannot corrupt your text
How we cut Yorùbá diacritic error by 58%, by changing the shape of the problem, and by noticing what our training corpus was quietly teaching us to do wrong.
Introducing DiacNet-1.1: Multilingual Diacritic Restoration for 10 Languages
Overview Diacritics-such as accents, tone marks, and special characters-are essential for clarity, pronunciation, and meaning in many languages. Yet, they are often stripped away in digital text due to technical limitations, user input habits, or data processing errors. This loss can lead to ambiguity, mispronunciation, and even miscommunication. Today, we’re excited to introduce DiacNet-1.1, a [...]
Pretrained From Scratch for Naija: Introducing mist-encoder-base-ng
mist-encoder-base-ng is a 30.9M encoder pretrained from scratch for Hausa, Yoruba, Igbo & Nigerian Pidgin, with a tokenizer that beats GPT-4o's on all four.
The Small Model Behind a Better Chat List: Introducing mist-tg-0.3b
mist-tg-0.3b is a free, open-source 300M model that turns a user's first chat message into a clean title. Runs on CPU, works across Latin-script languages.
One Passage In, Real Questions Out: Introducing mist-qg-1.5b
A compact, open-source question generator for 25 languages, and a data factory for building better retrieval systems. Every retrieval system, quiz app, and search engine shares a hidden hunger: it needs questions. Questions to train retrievers, questions to test rerankers, questions to turn static content into interactive learning. For English, that data exists in [...]
Know What Language It Is: Introducing the Olaverse LID Collection
Fast, open-source language identification built for African languages. If you've ever built a chatbot, moderation pipeline, or translation feature for Nigerian users, you've hit the same wall we did: before you can process text, you need to know what language it's in, and most off-the-shelf language detectors fall apart the moment they meet Yoruba [...]
Introducing diacnet-1.0: One Model, Ten Languages, and an Honest Look at What It Can (and Can't) Do Yet
Diacritics carry real meaning. In Yorùbá, the same base letters can spell three different words depending on the tone marks: ogun (war), ògùn (medicine), ogún (twenty). In Vietnamese, Polish, and Turkish, dropping accents doesn't just look wrong, it can make text ambiguous or outright unreadable. Yet diacritics are exactly what gets lost the moment [...]