Olaverse is an open-source multilingual AI infrastructure toolkit for building NLP, speech, retrieval, and language systems for underrepresented languages — built one focused model at a time.
608 languages identified · Yorùbá · Igbo · Hausa · Arabic · Vietnamese · Turkish · Polish — and beyond
Identify the language, restore its marks, generate questions from it. Benchmarked, maintained, and the models we'd point anyone to today.
Open weights, straight from Hugging Face, with the tools you already use. No API keys, no sign-ups, no rate limits.
from transformers import pipeline, AutoTokenizer, AutoModelForSeq2SeqLM # 1 — Which of 608 languages is this? lid = pipeline("text-classification", model="olaverse/lid-neural-608") lid("Ẹ kú àárọ̀, ṣé dáadáa ni?") >> {'label': 'yor_Latn', 'score': 0.99999...} # 2 — Restore the tone marks on unmarked text tok = AutoTokenizer.from_pretrained("olaverse/diacnet-2.0") model = AutoModelForSeq2SeqLM.from_pretrained("olaverse/diacnet-2.0") out = model.generate(**tok("<yor> O so fun ara re pe oun ko ni isoro kankan.", return_tensors="pt")) tok.decode(out[0], skip_special_tokens=True) >> "Ó sọ fún ara rẹ̀ pé òun kò ní ìṣòro kankan."
LATESTLanguage identification sounds solved. It is the first step of almost every multilingual pipeline: before youcan translate, filter a web crawl, route a support ticket or pick a speech model, you need to know what languagethe text is in. Open models like GlotLID and Meta's fastText LID cover hundreds of languages, and on long,clean paragraphs [...]



Release notes land on Hugging Face first.
Follow olaverse to catch every drop.
Every model is open-weight and free forever. Fine-tune it, deploy it, or drop it straight into your product. If you build something for underrepresented languages, tell us about it.