Restore Arabic tashkīl, with or without case endings, for text-to-speech, teaching material and readable copy, from models that also cover ten other languages.
Most Arabic is written without short vowels. Fluent readers fill them in from context; a text-to-speech engine, a language learner, or a screen reader cannot. The same consonant skeleton can be read several ways, and the wrong reading changes the meaning.
Dedicated Arabic diacritizers exist, but the generative ones can also alter the letters they were asked to vowel, and most are trained on Classical text, which is not what modern products publish.
Full diacritization includes the grammatical case ending (iʿrāb) on each word’s last letter, for teaching, classical and religious text, and TTS. No-case-endings mode leaves word-final vowels off and keeps shadda, which is how most modern Arabic is vowelled when it is vowelled at all.
diactag-2.0diactag-2.0 ranks second of five open Arabic diacritizers on Modern Standard Arabic benchmarks, and unlike the generative ones it never alters the input: compliance 1.0000. DER 0.0858 on WikiNews, 0.0669 without case endings.
diactag-2.0diacnet-2.0 handles Arabic with <ara> and <ara-nocase> from the same model it uses for the other ten languages: DER 0.095 on Classical Arabic, 0.072 without case endings. Keep alignment on: raw output changes a letter in some MSA sentences, and alignment halves the WikiNews error.
diacnet-2.0The Olaverse SDK wraps the models with sane defaults. If you would rather not add a dependency, the second tab is the same pipeline in plain transformers.
from olaverse.nlp import Diacritizer
ar = Diacritizer(model="diactag-2.0", lang="ar") # "ar" or "ara"
ar.restore("ذهب الطالب إلى المدرسة في الصباح")
# → 'ذَهَبَ الطَّالِبُ إلَى الْمَدْرَسَةِ فِي الصَّبَاحِ' full, for TTS and teaching
ar.restore("ذهب الطالب إلى المدرسة في الصباح", case_endings=False)
# → 'ذَهَب الطَّالِب إلَى الْمَدْرَسَة فِي الصَّبَاح' modern partial vowelling
# Generative alternative, aligned by default
Diacritizer(model="diacnet-2.0").restore("وهذا قول مرغوب عنه .", lang="ara")
# → 'وَهَذَا قَوْلٌ مَرْغُوبٌ عَنْهُ .'Every figure on this page comes from the published model cards. Download the weights and check us.
Restore diacritics and tone marks on text that was typed, scraped, or keyed in without them, across 11 languages from Yorùbá and Igbo to Vietnamese, Polish and Arabic.
/ Search & retrievalRestore the marks your documents lost, then match on an accent-folded key, so a user typing "oko" finds "ọkọ" and still sees it written properly.
The weights are open, so you can build it yourself this afternoon. If you would rather we tuned it to your domain, evaluated it properly, and handed it over, tell us what you are working on.