PRODUCTION MODELS · OPEN WEIGHTS ON HUGGING FACE

    Small, task-specific open models

    Olaverse is an open-source multilingual AI infrastructure toolkit for building NLP, speech, retrieval, and language systems for underrepresented languages — built one focused model at a time.

    🤗 Hugging Face

    608 languages identified · Yorùbá · Igbo · Hausa · Arabic · Vietnamese · Turkish · Polish — and beyond

    0
    LANGUAGES IDENTIFIED
    0
    LANGUAGES DIACRITIZED
    0
    LANGUAGES FOR QUESTION GEN

    Production models

    / READY FOR PRODUCTION

    Identify the language, restore its marks, generate questions from it. Benchmarked, maintained, and the models we'd point anyone to today.

    LID-608

    / Language identification

    DiacNet

    / Diacritic restoration

    MIST QG

    / Question generation

    Start in seconds

    Open weights, straight from Hugging Face, with the tools you already use. No API keys, no sign-ups, no rate limits.

    $ pip install olaverse
    quickstart.py
    from transformers import pipeline, AutoTokenizer, AutoModelForSeq2SeqLM
    
    # 1 — Which of 608 languages is this?
    lid = pipeline("text-classification", model="olaverse/lid-neural-608")
    lid("Ẹ kú àárọ̀, ṣé dáadáa ni?")
    >> {'label': 'yor_Latn', 'score': 0.99999...}
    
    # 2 — Restore the tone marks on unmarked text
    tok = AutoTokenizer.from_pretrained("olaverse/diacnet-2.0")
    model = AutoModelForSeq2SeqLM.from_pretrained("olaverse/diacnet-2.0")
    out = model.generate(**tok("<yor> O so fun ara re pe oun ko ni isoro kankan.", return_tensors="pt"))
    tok.decode(out[0], skip_special_tokens=True)
    >> "Ó sọ fún ara rẹ̀ pé òun kò ní ìṣòro kankan."
    / INSIGHTS & NEWS

    From the Blog

    608 languages in 37 MB: building African-first language identification
    LATEST
    Models

    608 languages in 37 MB: building African-first language identification

    Language identification sounds solved. It is the first step of almost every multilingual pipeline: before youcan translate, filter a web crawl, route a support ticket or pick a speech model, you need to know what languagethe text is in. Open models like GlotLID and Meta's fastText LID cover hundreds of languages, and on long,clean paragraphs [...]

    Olumide Ola·Oct 06, 2026Read
    diacnet-2.0: how we fixed Yorùbá, added Arabic, and removed 90% of our Hausa errors with one function
    Models

    diacnet-2.0: how we fixed Yorùbá, added Arabic, and removed 90% of our Hausa errors with one function

    Oct 06, 2026
    diactag-2.0: Arabic vowel marks from a model that cannot corrupt your text
    Models

    diactag-2.0: Arabic vowel marks from a model that cannot corrupt your text

    Oct 06, 2026
    A diacritic model that cannot corrupt your text
    Models

    A diacritic model that cannot corrupt your text

    Aug 04, 2026
    🤗

    Release notes land on Hugging Face first.
    Follow olaverse to catch every drop.

    Build with our models

    Every model is open-weight and free forever. Fine-tune it, deploy it, or drop it straight into your product. If you build something for underrepresented languages, tell us about it.