diacnet-2.0: how we fixed Yorùbá, added Arabic, and removed 90% of our Hausa errors with one function

diacnet restores diacritics by rewriting text: a byte-level
ByT5 model reads <yor> se eranko naa si gbo o? and writes ṣé ẹranko náà sì gbọ́ ọ?. Today we are releasing
diacnet-2.0 (582M) and
diacnet-mini-2.0 (300M): 11 languages including Arabic,
Yorùbá error cut by 60%, and, with one small post-processing step, lower error than our character tagger
diactag-2.0 on 8 of 10 languages.
This post is mostly about the three things that went wrong on the way, because each one taught us something.
1. More data made Yorùbá worse
diacnet-1.1 was trained on far more Yorùbá text than 1.0, and it got worse at Yorùbá: its diacritic error rate
rose from 0.155 to 0.201. The cause was measurable. Yorùbá is written with tone marks on almost every syllable,
but most Yorùbá on the web leaves them out. 1.0 had trained on about 2,000 carefully marked passages
(average mark density 0.565); 1.1’s much larger web corpus averaged 0.223. The model learned what it saw:
leave the tones off.
For 2.0 we attacked the data, not the model:
•Filter the weak rows. Yorùbá training pairs were kept only if their mark density was at least 0.6× that of well-marked text.
•Teacher-marked text. We took real Yorùbá, Igbo, Hausa and Arabic news and web text and had a large language model, DeepSeek-V4-Pro, re-mark it in full. A sentence was kept only if the teacher changed marks and nothing else. Where it lightly normalised a sentence (splitting a contraction, say), its version became the target and the input was rebuilt from it, so input and target always match letter for letter.
•Meaning-dependent words. We generated sentences built around words whose marks depend on meaning (Vietnamese ranh: rảnh “free”, rành “skilled”; Yorùbá ogun: ogun “war”, ogún “twenty”), with an English gloss for each sense.
The synthetic set came to about 37k pairs. Before choosing the teacher we
compared several on the same sentences; the gap between two strong teachers was itself informative, at 7.4% of
Yorùbá letters, 3.4% Igbo and 1.0% Hausa, a useful reminder that even the references a model learns from are
not perfectly agreed upon.
2. Our first base model barely learned anything new
We trained the small model first, from google/byt5-small, and it worked. For the larger model we tried to save
compute by continuing from diacnet-1.1 with a low learning rate (5e-5). It improved Yorùbá a little and
learned almost no Arabic: a 0.34 error rate, three times worse than the small model.
The fix was to stop being clever: we retrained base from google/byt5-base with the small model’s recipe
(learning rate 2e-4, 2.5M real pairs plus the teacher set, one epoch). The retrained base is now the best model
in the family on almost every language. Continued training at a cautious learning rate can preserve a model
so well that it cannot learn a new script.
3. The model was “wrong” because it was fixing typos
When we scored diacnet the same strict way we score diactag, Igbo and Hausa looked terrible: about 0.09 and
0.05 error rate against diactag’s 0.016 and 0.004. Reading the outputs showed why. Most errors were not wrong
marks. In 4-7% of sentences diacnet had changed a letter, and strict scoring counts every markable letter in
such a sentence as wrong.
This was a feature misfiring. diacnet’s training inputs include deliberate typos (7% character noise), so it
learned to fix spelling as well as add marks. On a benchmark, and in many real pipelines, that is unwanted.
The fix is a tiny function that runs after generation: line the output up with the input, keep the model’s
marks on every letter it left unchanged, and fall back to the input letter anywhere it changed, dropped or
added one. The result is guaranteed to be your text with marks added.
import difflib, unicodedata as ud
_FOLD = str.maketrans("ɓɗƙƴđıłƁƊƘƳĐŁ", "bdkydilBDKYDL") # letters with no combining form
_LETTER = {"ٓ", "ٔ", "ٕ"} # Arabic madda / hamza are spelling
def _units(text):
units = []
for c in ud.normalise("NFD", text):
if units and ud.combining(c):
if c in _LETTER:
units[-1][0] += c
units[-1][1] += c
else:
units.append([c, c])
return [(ud.normalise("NFC", b).translate(_FOLD), ud.normalise("NFC", f)) for b, f in units]
def align(source, output):
src, out = _units(source), _units(output)
res = [f for _, f in src]
sm = difflib.SequenceMatcher(None, [b for b, _ in src], [b for b, _ in out], autojunk=False)
for a, b, n in sm.get_matching_blocks():
res[a:a + n] = [f for _, f in out[b:b + n]]
return "".join(res)

| diacnet-2.0 | Raw | Aligned |
|---|---|---|
| Igbo | 0.093 | 0.018 |
| Hausa | 0.048 | 0.004 |
| Yorùbá | 0.078 | 0.070 |
| Arabic news (WikiNews) | 0.344 | 0.154 |
Hausa error falls by 90%, Igbo by 80%, and Arabic news by more than half, at no training cost. Users who do want
typo correction can still ask for the raw output.
Results
Scored with exactly the code and letter definitions used for diactag-2.0, on 1,000 sentences per language
from diacbench:

| diacnet-1.1 | diacnet-mini-2.0 | diacnet-2.0 | diactag-2.0 | |
|---|---|---|---|---|
| Yorùbá | 0.1743 | 0.0815 | 0.0698 | 0.0787 |
| Igbo | 0.0219 | 0.0202 | 0.0185 | 0.0162 |
| Hausa | 0.0063 | 0.0044 | 0.0041 | 0.0044 |
| Vietnamese | 0.0316 | 0.0207 | 0.0143 | 0.0141 |
| Polish | 0.0055 | 0.0024 | 0.0014 | 0.0035 |
| Spanish | 0.0064 | 0.0031 | 0.0027 | 0.0038 |
| French | 0.0028 | 0.0013 | 0.0012 | 0.0017 |
| Mean, 10 languages | 0.0261 | 0.0141 | 0.0117 | 0.0130 |
(Aligned output, diacbench v1 references; full tables with Turkish, Portuguese and Italian on the model cards.)
The two approaches now complement each other. diacnet-2.0 leads on Yorùbá and the European languages and
offers meaning hints and typo correction; diactag-2.0 leads on Igbo, Vietnamese and modern standard Arabic, is
15 times smaller, and runs in milliseconds on a CPU.
Arabic is the newest and weakest language for diacnet: error 0.096 on Classical Arabic and 0.154 on modern
news, behind diactag-2.0. Unlike most open Arabic diacritizers, diacnet was not trained on the Tashkeela
corpus that the Classical test sets come from, and it offers the same two modes: full, and without
grammatical case endings.
Meaning hints: a nudge, not a switch
You can tell diacnet what an ambiguous word means:
d.restore("Chi ay chi that su ranh vao nhung buoi toi sau khi da cho con ngu say.",
lang="vie", hints={"ranh": "free (time)"})
# 'Chị ấy chỉ thật sự rảnh vào những buổi tối sau khi đã cho con ngủ say.'
# without the hint: '... rành ...' ("skilled")
We tested this properly on 580 held-out sentences whose ambiguous spellings never appeared in training. Hints
lift Vietnamese accuracy on the ambiguous word from 0.63 to 0.90 and Turkish from 0.75 to 0.84. We also tried
the opposite: give the hint for the word’s other meaning and see whether the output follows it. Usually it
does not: diacnet switches between a quarter and a half of the time and otherwise trusts the sentence. Hints help when the
sentence leaves the meaning open; they will not override a sentence that makes it clear. We think that is the
right behaviour, but it is worth knowing.
Also in 2.0
•<auto>: leave the language out and the model infers it, at a cost of at most 0.0005 in error rate.
•Partial input: marks already in the text are kept (99.9-100% in our tests) and the rest are filled in, so Yorùbá typed with underdots but no tones comes back fully marked.
•Two sizes: diacnet-2.0 is more accurate everywhere (17% lower mean error); diacnet-mini-2.0 is half the size and within a hair of it on Igbo, Hausa, French, Spanish and Italian.
Try it
pip install "olaverse[deeplearning]"
from olaverse.nlp import Diacritizer
d = Diacritizer(model="diacnet-2.0", lang="yor", device="cuda") # or "cpu" / "mps"
d.restore("O so fun ara re pe oun ko ni isoro kankan.")
# 'Ó sọ fún ara rẹ̀ pé òun kò ní ìṣòro kankan.'
d.restore("وهذا قول مرغوب عنه .", lang="ara", case_endings=False)
# 'وَهَذَا قَوْل مَرْغُوب عَنْه .'
The library aligns the output by default (aligned=False returns the raw text) and handles chunking and
batching. The model cards also show plain transformers usage with the alignment function.
•Models: olaverse/diacnet-2.0, olaverse/diacnet-mini-2.0
•Character tagger: olaverse/diactag-2.0
•Benchmark: olaverse/diacbench
•Library: olaverse
Enjoyed this? Every model we write about is open weight and free to download.


