/ Insights & news

    Latest thinking

    Insights on artificial intelligence, software engineering, and the future of digital product development.

    Models Oct 06, 2026

    608 languages in 37 MB: building African-first language identification

    Language identification sounds solved. It is the first step of almost every multilingual pipeline: before youcan translate, filter a web crawl, route a support ticket or pick a speech model, you need to know what languagethe text is in. Open models like GlotLID and Meta's fastText LID cover hundreds of languages, and on long,clean paragraphs [...]

    Read article
    Models Oct 06, 2026

    diacnet-2.0: how we fixed Yorùbá, added Arabic, and removed 90% of our Hausa errors with one function

    diacnet restores diacritics by rewriting text: a byte-levelByT5 model reads <yor> se eranko naa si gbo o? and writes ṣé ẹranko náà sì gbọ́ ọ?. Today we are releasingdiacnet-2.0 (582M) anddiacnet-mini-2.0 (300M): 11 languages including Arabic,Yorùbá error cut by 60%, and, with one small post-processing step, lower error than our character taggerdiactag-2.0 on 8 of 10 languages. This post is mostly about [...]

    Read article
    Models Oct 06, 2026

    diactag-2.0: Arabic vowel marks from a model that cannot corrupt your text

    Diacritics carry meaning. In Yorùbá, ọkọ́ is a hoe, ọkọ̀ a vehicle and ọkọ a husband. In Arabic, theshort vowels that distinguish kataba (he wrote) from kutiba (it was written) are usually not written atall. Text without its marks is ambiguous for readers, for text-to-speech, and for every model downstream. diactag-1.0 restored diacritics in ten languages with an unusualguarantee: it cannot change your text. It never generates [...]

    Read article
    Models Aug 04, 2026

    A diacritic model that cannot corrupt your text

    How we cut Yorùbá diacritic error by 58%, by changing the shape of the problem, and by noticing what our training corpus was quietly teaching us to do wrong.

    Read article
    Models Jul 31, 2026

    Introducing DiacNet-1.1: Multilingual Diacritic Restoration for 10 Languages

    Overview Diacritics-such as accents, tone marks, and special characters-are essential for clarity, pronunciation, and meaning in many languages. Yet, they are often stripped away in digital text due to technical limitations, user input habits, or data processing errors. This loss can lead to ambiguity, mispronunciation, and even miscommunication. Today, we’re excited to introduce DiacNet-1.1, a [...]

    Read article
    Models Jul 19, 2026

    Pretrained From Scratch for Naija: Introducing mist-encoder-base-ng

    mist-encoder-base-ng is a 30.9M encoder pretrained from scratch for Hausa, Yoruba, Igbo & Nigerian Pidgin, with a tokenizer that beats GPT-4o's on all four.

    Read article
    Models Jul 18, 2026

    The Small Model Behind a Better Chat List: Introducing mist-tg-0.3b

    mist-tg-0.3b is a free, open-source 300M model that turns a user's first chat message into a clean title. Runs on CPU, works across Latin-script languages.

    Read article
    Models Jul 16, 2026

    One Passage In, Real Questions Out: Introducing mist-qg-1.5b

    A compact, open-source question generator for 25 languages, and a data factory for building better retrieval systems. Every retrieval system, quiz app, and search engine shares a hidden hunger: it needs questions. Questions to train retrievers, questions to test rerankers, questions to turn static content into interactive learning. For English, that data exists in [...]

    Read article
    Models Jul 16, 2026

    Know What Language It Is: Introducing the Olaverse LID Collection

    Fast, open-source language identification built for African languages. If you've ever built a chatbot, moderation pipeline, or translation feature for Nigerian users, you've hit the same wall we did: before you can process text, you need to know what language it's in, and most off-the-shelf language detectors fall apart the moment they meet Yoruba [...]

    Read article
    Models Jul 14, 2026

    Introducing diacnet-1.0: One Model, Ten Languages, and an Honest Look at What It Can (and Can't) Do Yet

    Diacritics carry real meaning. In Yorùbá, the same base letters can spell three different words depending on the tone marks: ogun (war), ògùn (medicine), ogún (twenty). In Vietnamese, Polish, and Turkish, dropping accents doesn't just look wrong, it can make text ambiguous or outright unreadable. Yet diacritics are exactly what gets lost the moment [...]

    Read article