diactag-2.0: Arabic vowel marks from a model that cannot corrupt your text

Diacritics carry meaning. In Yorùbá, ọkọ́ is a hoe, ọkọ̀ a vehicle and ọkọ a husband. In Arabic, the
short vowels that distinguish kataba (he wrote) from kutiba (it was written) are usually not written at
all. Text without its marks is ambiguous for readers, for text-to-speech, and for every model downstream.
diactag-1.0 restored diacritics in ten languages with an unusual
guarantee: it cannot change your text. It never generates characters. It reads the input and, for each
letter, classifies which marks belong on it, so the output is always the input with marks added. That matters
more than it sounds: when we asked Claude Sonnet 4.5 only to add marks, it rewrote 26% of Yorùbá and 34% of
Hausa sentences (from the diacnet-1.1 evaluation). The
diactag-1.0 write-up covers
the design in detail.
diactag-2.0 adds Arabic, keeps the guarantee, and keeps
1.0’s accuracy on the original ten languages: Yorùbá, Igbo, Hausa, Vietnamese, Polish, Turkish, Portuguese,
Spanish, French and Italian. It is 37.9M parameters, ships a 38.6 MB int8 ONNX export, and restores a typical
sentence in about 8 ms on a laptop CPU.
Fitting Arabic into a tagger built for Latin script
diactag factors every character into a base letter plus two labels: a SHAPE (letter identity, such as ẹ vs e,
ơ vs o, ł vs l) and a TONE (tone, stress or accent). Arabic maps onto that cleanly once you decide what is
spelling and what is a mark:
•Vowels, tanwīn and sukūn (fatḥa, ḍamma, kasra, -an/-un/-in, sukūn) become TONE classes.
•Shadda (consonant doubling) and dagger alif become SHAPE classes, so a letter can carry a shadda and a vowel at the same time.
•Hamza letters (أ إ آ ؤ ئ) are spelling, not diacritics. Unicode would normally decompose أ into alif plus a combining hamza; diactag-2.0 deliberately does not, so the model never adds or removes a hamza. If the input spells a hamza, it stays; if it doesn’t, the model won’t invent one.
•A legality mask per language rules out impossible combinations, such as a Yorùbá tone on an Arabic letter or an Arabic vowel on a Latin one.
The label space grew from spec 1.x to 2.0.0: 986 characters, 18 shape classes and 15 tone classes. The
important detail is that every existing label kept its index. Instead of training from scratch, we copied
diactag-1.0’s weights into the larger output layer unchanged, added fresh rows for Arabic, and continued
training for 20,000 steps, about two hours on one A100. Measured side by side with the same scoring code, 2.0
matches 1.0 on the original languages (mean diacritic error rate 0.0130 against 0.0127), so adding a language
cost nothing on the existing ones.
Two kinds of Arabic vowelling
Fully vowelled Arabic includes the grammatical case ending on almost every word’s last letter: it is required
for teaching material, Qur’anic and classical text, and text-to-speech. But it is also the hardest part of
the task, because it depends on syntax, and most modern Arabic that is vowelled at all leaves it off.
diactag-2.0 does both from one model:
from olaverse.nlp import Diacritizer
d = Diacritizer(model="diactag-2.0", lang="ar")
d.restore("ذهب الطالب إلى المدرسة في الصباح")
# 'ذَهَبَ الطَّالِبُ إلَى الْمَدْرَسَةِ فِي الصَّبَاحِ'
d.restore("ذهب الطالب إلى المدرسة في الصباح", case_endings=False)
# 'ذَهَب الطَّالِب إلَى الْمَدْرَسَة فِي الصَّبَاح'
case_endings=False drops the vowel, tanwīn or sukūn from each word’s final letter and keeps the shadda. On
Classical Arabic it cuts the error rate from 0.059 to 0.038.
How it compares
We scored diactag-2.0 against four open Arabic diacritizers on the same sentences, counting Arabic letters
only. Scoring is strict: an output that changes, adds or drops a letter counts the whole sentence as wrong,
because in a real pipeline it is wrong.

| System | Type | WikiNews (modern news) | SadeedDiac-25 benchmark | Sentences with altered letters |
|---|---|---|---|---|
| CATT-EO | character transformer | 0.040 | 0.047 | 0.1% |
| diactag-2.0 | character tagger, 37.9M | 0.086 | 0.099 | 0.0% |
| Shakkala v3 | BiLSTM | 0.106 | 0.107 | 0.0% |
| Fine-Tashkeel | ByT5, generative | 0.149 | 0.216 | 14.2% |
| Tashkeel-700M | LLM, generative | 0.296 | 0.309 | 24.4% |
On modern standard Arabic, diactag-2.0 ranks second of the five, behind CATT, at a fraction of the size of the
generative systems, and it is one of the two that never touch the input. The two generative models alter the
letters of one sentence in seven and one in four respectively, which is exactly the failure diactag was built
to rule out. Modern Arabic is where diactag-2.0 has the most room to grow.
We audited our own benchmark
Releasing diactag-2.0 meant looking hard at diacbench,
the benchmark we use for every diacritics model. Reading model errors closely, we were no longer sure every
“error” was the model’s: some looked like errors in the references, such as missing underdots and wrong
tones in Yorùbá, missing dotted vowels in Igbo, and Hausa kasar where the correct spelling is ƙasar.
So we ran an audit:
1.Two second opinions. DeepSeek-V4-Pro (a large LLM) and diactag-2.0 (a tagger) re-marked every sentence independently.
2.Flags. A word was flagged only when both systems agreed with each other and disagreed with the reference: 1,537 flags, mostly Yorùbá (6.9% of words), Igbo (2.4%) and Hausa (1.1%).
3.An independent judge. A third model, Claude Sonnet 5.5, saw each flagged word in its sentence and chose the reference form, the flagged form, both or neither, with a confidence. Only high-confidence choices of the flagged form became corrections: 609 of the 1,537.
4.A manual check. On 30 Yorùbá flags, all 12 accepted corrections matched a manual judgement; across about 40 accepted corrections checked in five languages, one was wrong and was reverted.
diacbench v2 is public, with every change listed in a changelog and the original v1 references kept in their
own column. On v2, diactag-2.0’s Yorùbá error falls from 0.079 to 0.073 and Hausa from 0.0044 to 0.0025.
There is a catch, and the dataset card says it plainly: corrections were only proposed where the two
re-marking systems agreed, so v2 cannot fix errors they share, and its gains favour systems that behave like
them. That is why we report v1 and v2 side by side, and why we recommend everyone does.
Try it
pip install "olaverse[deeplearning]" # add [onnx] for the int8 ONNX backend
from olaverse.nlp import Diacritizer
d = Diacritizer(model="diactag-2.0", lang="yor")
d.restore("se eranko naa si gbo o?") # 'ṣé ẹranko náà sì gbọ́ ọ?'
Diacritizer(model="diactag-2.0").restore("Toi khong biet tieng Viet") # language detected: 'Tôi không biết tiếng Việt'
fast = Diacritizer(model="diactag-2.0", lang="yor", onnx=True) # int8 ONNX on CPU
text, details = d.restore("se eranko naa si gbo o?", return_details=True) # per-character confidence
The full training and inference code, the training notebook and the benchmark notebook are in the
model repository, so it can be reproduced or fine-tuned.
If you need the highest accuracy on Yorùbá or the European languages and can run a larger model, see
diacnet-2.0, our generative diacritizer released alongside it.
•Model: olaverse/diactag-2.0
•Benchmark: olaverse/diacbench
•Library: olaverse
Enjoyed this? Every model we write about is open weight and free to download.


