Models Oct 06, 2026 5 min read

    diactag-2.0: Arabic vowel marks from a model that cannot corrupt your text

    OO
    Olumide Ola
    Olaverse Lab
    diactag-2.0: Arabic vowel marks from a model that cannot corrupt your text

    Diacritics carry meaning. In Yorùbá, ọkọ́ is a hoe, ọkọ̀ a vehicle and ọkọ a husband. In Arabic, the
    short vowels that distinguish kataba (he wrote) from kutiba (it was written) are usually not written at
    all. Text without its marks is ambiguous for readers, for text-to-speech, and for every model downstream.

    diactag-1.0 restored diacritics in ten languages with an unusual
    guarantee: it cannot change your text. It never generates characters. It reads the input and, for each
    letter, classifies which marks belong on it, so the output is always the input with marks added. That matters
    more than it sounds: when we asked Claude Sonnet 4.5 only to add marks, it rewrote 26% of Yorùbá and 34% of
    Hausa sentences (from the diacnet-1.1 evaluation). The
    diactag-1.0 write-up covers
    the design in detail.

    diactag-2.0 adds Arabic, keeps the guarantee, and keeps
    1.0’s accuracy on the original ten languages: Yorùbá, Igbo, Hausa, Vietnamese, Polish, Turkish, Portuguese,
    Spanish, French and Italian. It is 37.9M parameters, ships a 38.6 MB int8 ONNX export, and restores a typical
    sentence in about 8 ms on a laptop CPU.

    Fitting Arabic into a tagger built for Latin script

    diactag factors every character into a base letter plus two labels: a SHAPE (letter identity, such as ẹ vs e,
    ơ vs o, ł vs l) and a TONE (tone, stress or accent). Arabic maps onto that cleanly once you decide what is
    spelling and what is a mark:

    •Vowels, tanwīn and sukūn (fatḥa, ḍamma, kasra, -an/-un/-in, sukūn) become TONE classes.

    •Shadda (consonant doubling) and dagger alif become SHAPE classes, so a letter can carry a shadda and a vowel at the same time.

    •Hamza letters (أ إ آ ؤ ئ) are spelling, not diacritics. Unicode would normally decompose أ into alif plus a combining hamza; diactag-2.0 deliberately does not, so the model never adds or removes a hamza. If the input spells a hamza, it stays; if it doesn’t, the model won’t invent one.

    •A legality mask per language rules out impossible combinations, such as a Yorùbá tone on an Arabic letter or an Arabic vowel on a Latin one.

    The label space grew from spec 1.x to 2.0.0: 986 characters, 18 shape classes and 15 tone classes. The
    important detail is that every existing label kept its index. Instead of training from scratch, we copied
    diactag-1.0’s weights into the larger output layer unchanged, added fresh rows for Arabic, and continued
    training for 20,000 steps, about two hours on one A100. Measured side by side with the same scoring code, 2.0
    matches 1.0 on the original languages (mean diacritic error rate 0.0130 against 0.0127), so adding a language
    cost nothing on the existing ones.

    Two kinds of Arabic vowelling

    Fully vowelled Arabic includes the grammatical case ending on almost every word’s last letter: it is required
    for teaching material, Qur’anic and classical text, and text-to-speech. But it is also the hardest part of
    the task, because it depends on syntax, and most modern Arabic that is vowelled at all leaves it off.

    diactag-2.0 does both from one model:

    from olaverse.nlp import Diacritizer
    
    d = Diacritizer(model="diactag-2.0", lang="ar")
    d.restore("ذهب الطالب إلى المدرسة في الصباح")
    # 'ذَهَبَ الطَّالِبُ إلَى الْمَدْرَسَةِ فِي الصَّبَاحِ'
    d.restore("ذهب الطالب إلى المدرسة في الصباح", case_endings=False)
    # 'ذَهَب الطَّالِب إلَى الْمَدْرَسَة فِي الصَّبَاح'

    case_endings=False drops the vowel, tanwīn or sukūn from each word’s final letter and keeps the shadda. On
    Classical Arabic it cuts the error rate from 0.059 to 0.038.

    How it compares

    We scored diactag-2.0 against four open Arabic diacritizers on the same sentences, counting Arabic letters
    only. Scoring is strict: an output that changes, adds or drops a letter counts the whole sentence as wrong,
    because in a real pipeline it is wrong.

    Arabic: diactag-2.0 vs open Arabic diacritizers
    SystemTypeWikiNews (modern news)SadeedDiac-25 benchmarkSentences with altered letters
    CATT-EOcharacter transformer0.0400.0470.1%
    diactag-2.0character tagger, 37.9M0.0860.0990.0%
    Shakkala v3BiLSTM0.1060.1070.0%
    Fine-TashkeelByT5, generative0.1490.21614.2%
    Tashkeel-700MLLM, generative0.2960.30924.4%

    On modern standard Arabic, diactag-2.0 ranks second of the five, behind CATT, at a fraction of the size of the
    generative systems, and it is one of the two that never touch the input. The two generative models alter the
    letters of one sentence in seven and one in four respectively, which is exactly the failure diactag was built
    to rule out. Modern Arabic is where diactag-2.0 has the most room to grow.

    We audited our own benchmark

    Releasing diactag-2.0 meant looking hard at diacbench,
    the benchmark we use for every diacritics model. Reading model errors closely, we were no longer sure every
    “error” was the model’s: some looked like errors in the references, such as missing underdots and wrong
    tones in Yorùbá, missing dotted vowels in Igbo, and Hausa kasar where the correct spelling is ƙasar.

    So we ran an audit:

    1.Two second opinions. DeepSeek-V4-Pro (a large LLM) and diactag-2.0 (a tagger) re-marked every sentence independently.

    2.Flags. A word was flagged only when both systems agreed with each other and disagreed with the reference: 1,537 flags, mostly Yorùbá (6.9% of words), Igbo (2.4%) and Hausa (1.1%).

    3.An independent judge. A third model, Claude Sonnet 5.5, saw each flagged word in its sentence and chose the reference form, the flagged form, both or neither, with a confidence. Only high-confidence choices of the flagged form became corrections: 609 of the 1,537.

    4.A manual check. On 30 Yorùbá flags, all 12 accepted corrections matched a manual judgement; across about 40 accepted corrections checked in five languages, one was wrong and was reverted.

    diacbench v2 is public, with every change listed in a changelog and the original v1 references kept in their
    own column. On v2, diactag-2.0’s Yorùbá error falls from 0.079 to 0.073 and Hausa from 0.0044 to 0.0025.

    There is a catch, and the dataset card says it plainly: corrections were only proposed where the two
    re-marking systems agreed, so v2 cannot fix errors they share, and its gains favour systems that behave like
    them. That is why we report v1 and v2 side by side, and why we recommend everyone does.

    Try it

    pip install "olaverse[deeplearning]"     # add [onnx] for the int8 ONNX backend
    from olaverse.nlp import Diacritizer
    
    d = Diacritizer(model="diactag-2.0", lang="yor")
    d.restore("se eranko naa si gbo o?")                     # 'ṣé ẹranko náà sì gbọ́ ọ?'
    
    Diacritizer(model="diactag-2.0").restore("Toi khong biet tieng Viet")   # language detected: 'Tôi không biết tiếng Việt'
    
    fast = Diacritizer(model="diactag-2.0", lang="yor", onnx=True)          # int8 ONNX on CPU
    text, details = d.restore("se eranko naa si gbo o?", return_details=True)   # per-character confidence

    The full training and inference code, the training notebook and the benchmark notebook are in the
    model repository, so it can be reproduced or fine-tuned.

    If you need the highest accuracy on Yorùbá or the European languages and can run a larger model, see
    diacnet-2.0, our generative diacritizer released alongside it.

    •Model: olaverse/diactag-2.0

    •Benchmark: olaverse/diacbench

    •Library: olaverse

    Enjoyed this? Every model we write about is open weight and free to download.

    🤗 Follow on Hugging Face

    More from the Blog

    608 languages in 37 MB: building African-first language identification
    Models

    608 languages in 37 MB: building African-first language identification

    diacnet-2.0: how we fixed Yorùbá, added Arabic, and removed 90% of our Hausa errors with one function
    Models

    diacnet-2.0: how we fixed Yorùbá, added Arabic, and removed 90% of our Hausa errors with one function

    A diacritic model that cannot corrupt your text
    Models

    A diacritic model that cannot corrupt your text