Add the vowel marks Arabic text leaves out

    Restore Arabic tashkīl, with or without case endings, for text-to-speech, teaching material and readable copy, from models that also cover ten other languages.

    Text-to-speechEdtech & Quran studyPublishingAccessibility
    0.0587
    DER, CLASSICAL ARABIC (DIACTAG-2.0)
    0.0382
    DER WITHOUT CASE ENDINGS
    1.0000
    COMPLIANCE, LETTERS NEVER CHANGE
    / The problem

    Why this breaks today

    Most Arabic is written without short vowels. Fluent readers fill them in from context; a text-to-speech engine, a language learner, or a screen reader cannot. The same consonant skeleton can be read several ways, and the wrong reading changes the meaning.

    Dedicated Arabic diacritizers exist, but the generative ones can also alter the letters they were asked to vowel, and most are trained on Classical text, which is not what modern products publish.

    / How it works

    The pipeline, step by step

    01

    Choose full marks or no case endings

    Full diacritization includes the grammatical case ending (iʿrāb) on each word’s last letter, for teaching, classical and religious text, and TTS. No-case-endings mode leaves word-final vowels off and keeps shadda, which is how most modern Arabic is vowelled when it is vowelled at all.

    diactag-2.0
    02

    Use the tagger for Modern Standard Arabic

    diactag-2.0 ranks second of five open Arabic diacritizers on Modern Standard Arabic benchmarks, and unlike the generative ones it never alters the input: compliance 1.0000. DER 0.0858 on WikiNews, 0.0669 without case endings.

    diactag-2.0
    03

    Or the generator, with alignment on

    diacnet-2.0 handles Arabic with <ara> and <ara-nocase> from the same model it uses for the other ten languages: DER 0.095 on Classical Arabic, 0.072 without case endings. Keep alignment on: raw output changes a letter in some MSA sentences, and alignment halves the WikiNews error.

    diacnet-2.0
    / The code

    Running in about ten lines

    The Olaverse SDK wraps the models with sane defaults. If you would rather not add a dependency, the second tab is the same pipeline in plain transformers.

    $pip install "olaverse[deeplearning]"
    tashkeel.py
    from olaverse.nlp import Diacritizer
    
    ar = Diacritizer(model="diactag-2.0", lang="ar")       # "ar" or "ara"
    
    ar.restore("ذهب الطالب إلى المدرسة في الصباح")
    # → 'ذَهَبَ الطَّالِبُ إلَى الْمَدْرَسَةِ فِي الصَّبَاحِ'      full, for TTS and teaching
    
    ar.restore("ذهب الطالب إلى المدرسة في الصباح", case_endings=False)
    # → 'ذَهَب الطَّالِب إلَى الْمَدْرَسَة فِي الصَّبَاح'        modern partial vowelling
    
    # Generative alternative, aligned by default
    Diacritizer(model="diacnet-2.0").restore("وهذا قول مرغوب عنه .", lang="ara")
    # → 'وَهَذَا قَوْلٌ مَرْغُوبٌ عَنْهُ .'
    / This is you if
    • You run Arabic text-to-speech and the voice misreads unvowelled words.
    • You build Arabic learning material and need fully vowelled text.
    • You need marks added without any risk of the letters themselves changing.

    Every figure on this page comes from the published model cards. Download the weights and check us.

    Want this on your data?

    The weights are open, so you can build it yourself this afternoon. If you would rather we tuned it to your domain, evaluated it properly, and handed it over, tell us what you are working on.