AI Models
    DiacNet Mini 2.0

    DiacNet Mini 2.0

    DiacNet family · open weights on Hugging Face.

    Yorùbá diacritic and tone restoration powered by a multilingual ByT5 model, making text readable, searchable, and machine-usable.

    TransformersSafetensorsT5Text2text GenerationDiacriticsDiacritizationAccent RestorationTone RestorationTashkeelYorubaIgboHausaVietnameseArabicByt5Text GenerationYorubaIgboHausaViPlTrPtEsFrItArDataset:olaverse/diacnet 1.1 TrainDataset:olaverse/diacbenchBase: google/byt5-smallBase: google/byt5-smallLicence: apache-2.0Text Generation InferenceEndpoints Compatible
    0
    Likes
    2
    Downloads
    Oct 2026
    Created
    / Model card

    Built for production use

    Open-weights repository on Hugging Face

    Integrated ecosystem protocol tier.

    Compatible with Transformers library

    Integrated ecosystem protocol tier.

    Optimised for low-latency inference

    Integrated ecosystem protocol tier.

    Community engagement: 0 likes

    Integrated ecosystem protocol tier.

    / Quick start

    Use DiacNet Mini 2.0, straight from its model card.

    example.pyfrom the model card
    import torch
    from transformers import AutoTokenizer, AutoModelForSeq2SeqLM
    repo = "olaverse/diacnet-mini-2.0"
    tok = AutoTokenizer.from_pretrained(repo)
    model = AutoModelForSeq2SeqLM.from_pretrained(repo, dtype=torch.bfloat16).eval() # float32 on CPU
    import difflib
    import unicodedata as ud
    _FOLD = str.maketrans("ɓɗƙƴđıłƁƊƘƳĐŁ", "bdkydilBDKYDL") # letters with no combining form
    _LETTER = {"\u0653", "\u0654", "\u0655"} # Arabic madda / hamza are spelling, not marks
    def _units(text):
    units = [] # [letter, letter + marks]
    for c in ud.normalize("NFD", text):
    if units and ud.combining(c):
    if c in _LETTER:
    units[-1][0] += c
    units[-1][1] += c
    else:
    units.append([c, c])
    return [(ud.normalize("NFC", b).translate(_FOLD), ud.normalize("NFC", f)) for b, f in units]
    def align(source, output):
    """`source` with the marks diacnet put on every letter it kept; letters it changed,
    dropped or added fall back to `source`, so your text itself never changes."""
    src, out = _units(source), _units(output)
    res = [f for _, f in src]
    sm = difflib.SequenceMatcher(None, [b for b, _ in src], [b for b, _ in out], autojunk=False)
    # ... see the full example on the model card
    / Model card

    Official repository README

    Dynamically loaded from Hugging Face

    Model card metadata is available on Hugging Face.

    / Built with
    Transformers PyTorch Python

    Ready to try DiacNet Mini 2.0?

    Yorùbá diacritic and tone restoration powered by a multilingual ByT5 model, making text readable, searchable, and machine-usable.

    / Model ecosystem

    Explore more models