/ Insights & news

    Latest thinking

    Insights on artificial intelligence, software engineering, and the future of digital product development.

    Models Aug 04, 2026

    A diacritic model that cannot corrupt your text

    How we cut Yorùbá diacritic error by 58%, by changing the shape of the problem, and by noticing what our training corpus was quietly teaching us to do wrong.

    Read article
    Models Jul 31, 2026

    Introducing DiacNet-1.1: Multilingual Diacritic Restoration for 10 Languages

    Overview Diacritics-such as accents, tone marks, and special characters-are essential for clarity, pronunciation, and meaning in many languages. Yet, they are often stripped away in digital text due to technical limitations, user input habits, or data processing errors. This loss can lead to ambiguity, mispronunciation, and even miscommunication. Today, we’re excited to introduce DiacNet-1.1, a [...]

    Read article
    Models Jul 19, 2026

    Pretrained From Scratch for Naija: Introducing mist-encoder-base-ng

    mist-encoder-base-ng is a 30.9M encoder pretrained from scratch for Hausa, Yoruba, Igbo & Nigerian Pidgin, with a tokenizer that beats GPT-4o's on all four.

    Read article
    Models Jul 18, 2026

    The Small Model Behind a Better Chat List: Introducing mist-tg-0.3b

    mist-tg-0.3b is a free, open-source 300M model that turns a user's first chat message into a clean title. Runs on CPU, works across Latin-script languages.

    Read article
    Models Jul 16, 2026

    One Passage In, Real Questions Out: Introducing mist-qg-1.5b

    A compact, open-source question generator for 25 languages, and a data factory for building better retrieval systems. Every retrieval system, quiz app, and search engine shares a hidden hunger: it needs questions. Questions to train retrievers, questions to test rerankers, questions to turn static content into interactive learning. For English, that data exists in [...]

    Read article
    Models Jul 16, 2026

    Know What Language It Is: Introducing the Olaverse LID Collection

    Fast, open-source language identification built for African languages. If you've ever built a chatbot, moderation pipeline, or translation feature for Nigerian users, you've hit the same wall we did: before you can process text, you need to know what language it's in, and most off-the-shelf language detectors fall apart the moment they meet Yoruba [...]

    Read article
    Models Jul 14, 2026

    Introducing diacnet-1.0: One Model, Ten Languages, and an Honest Look at What It Can (and Can't) Do Yet

    Diacritics carry real meaning. In Yorùbá, the same base letters can spell three different words depending on the tone marks: ogun (war), ògùn (medicine), ogún (twenty). In Vietnamese, Polish, and Turkish, dropping accents doesn't just look wrong, it can make text ambiguous or outright unreadable. Yet diacritics are exactly what gets lost the moment [...]

    Read article