The ground floor of the stack: otk-bpe-50k is trained for Yorùbá, Igbo, Hausa, and Nigerian Pidgin, so African-language text stops being tokenised into confetti. otk-bpe targets a different set of languages (Swahili, Kinyarwanda, French) — the two are parallel vocabularies, not versions of each other.


OTK TokenizersLABSByte-level BPE tokenizer built for Nigerian languages, Yorùbá, Igbo, Hausa, and Pidgin, efficient subword segmentation where general vocabularies fall apart.

OTK TokenizersLABSByte-level BPE tokenizer built for multilingual text, efficient subword segmentation where general vocabularies fall apart.
More of the stack, organised the same way.