PurpleMIST: decision models that speak eight African languages

Meet PurpleMIST-Flash: a 7.9B decision model that matches DeepSeek-V4-Pro, a model about 200 times its size,
across eight African languages. Its probabilities are about five times closer to the truth.
A decision model doesn’t write text. You give it a piece of state (a WhatsApp message, a support ticket, a
call transcript, a form record) and a set of typed questions: which team should handle this?, is the customer
asking for a refund?, how urgent is it? For every question it returns a probability for each option, in one
forward pass, in milliseconds. It is the fastest way to triage, route, moderate and automate at scale. With
TypeSafe’s Jev, more than a hundred open reproductions, and Ollama’s new /v1/systemone endpoint, it has become
one of the most active corners of open AI.
Until now, almost all of it worked only in English. Today we’re changing that with three releases:
•PurpleMIST-Flash-1.0 (7.9B): our flagship decision model, for English and eight African languages.
•PurpleMIST-Mini-1.0 (1.9B): the same interface in a model small enough for almost any GPU.
•African Typed Decisions: to our knowledge, the first decision benchmark for African languages. It has 3,739 cases and 16,881 decisions in Amharic, Hausa, Igbo, Nigerian Pidgin, Somali, Swahili, Yorùbá and isiZulu, and an open leaderboard.
Frontier accuracy at a fraction of the size
On African Typed Decisions, PurpleMIST-Flash scores 0.668 accuracy, level with DeepSeek-V4-Pro (0.669), a
1.6-trillion-parameter model. Where it pulls ahead is the part that matters for a decision model: how much you
can trust each probability.
| Model | Size | English | 8 African languages | KL ↓ | Brier ↓ |
|---|---|---|---|---|---|
| PurpleMIST-Flash-1.0 | 7.9B | 0.720 | 0.668 | 0.214 | 0.115 |
| DeepSeek-V4-Pro (prompted) | 1.6T | 0.720 | 0.669 | 1.125 | 0.199 |
| Qwen3.8-Flash (prompted) | undisclosed | 0.712 | 0.643 | 0.862 | 0.216 |
| OpenDecider-small | 4B | 0.648 | 0.477 | 0.365 | 0.199 |
| laya-multilingual | 0.32B | 0.348 | 0.322 | 5.053 | 0.545 |
Other open decision models collapse outside English: OpenDecider-small loses 0.171 accuracy. PurpleMIST-Flash
loses just 0.053. It leads every open decision model in every one of the eight languages.
Strong in English too
•LocalLLaMA/typed-decisions: 0.720 accuracy, zero-shot. Its probability quality, KL 0.189 and Brier 0.098, ranks 2nd among the published zero-shot models, ahead of Liquid AI d1 and TypeSafe Jev.
•Jev Decision Index (37 public benchmarks, run with the maintainers’ own kit): ahead of 8 of the 13 models of its size. It is 4.8 points above their median, and above the median in language understanding (+6.4), knowledge and reasoning (+4.3), tools (+3.8) and retrieval (+1.4). We’ve submitted it for the index’s private evaluation.
Probabilities you can act on
A decision model is only as useful as its confidence. When it says 90%, it should be right about 90% of the time.
Every PurpleMIST model is calibrated across domains: one temperature per question type, fitted on thousands of
states it never saw in training, drawn evenly from eight different sources. The result is an expected
calibration error of 0.049 across the eight African languages and 0.056 in English. So you can set a
threshold once and trust it: auto-approve above 0.9, send to a human below.
Fast at any length
•Speed: a median of 63 ms per request on a single GPU, over 140,000 benchmark requests of every size.
•Long inputs: up to 65,536 tokens per pass. Long contracts, transcripts and documents are read whole, never truncated.
•Many questions: any number of them about a state share one pass, and any number of options per question works, because the model reads each option’s text rather than memorising labels.
How it works
PurpleMIST-Flash is built on the text model of Qwen3.5-9B with a purpose-built pointer head. The whole state
and all its questions go in together, and each option is scored directly against the question it belongs to. It
was trained on 300,000 real and synthetic decision states, including 86,134 decisions written natively in, or
translated into, the eight African languages.
Start in minutes
import os, sys
from huggingface_hub import hf_hub_download
repo = "olaverse/PurpleMIST-Flash-1.0"
sys.path.insert(0, os.path.dirname(hf_hub_download(repo, "purplemist.py")))
from purplemist import Decider
d = Decider.from_pretrained(repo, device="cuda")
d.decide({"message": "Ẹ jọ̀ọ́, wọ́n ti yọ owó lẹ́ẹ̀mejì lórí káàdì mi. Mo fẹ́ kí ẹ dá owó mi padà."},
{"team": {"type": "choice", "instructions": "Which team should handle this?",
"criteria": {"billing": "payments and refunds", "delivery": "orders in transit"}},
"refund": {"type": "noul", "instructions": "The customer is asking for money back."}})
Or run it as a server. serve.py ships in each model repo and uses the same POST /v1/systemone request and
answer format as Ollama’s decision endpoint, so code written for that API reads PurpleMIST’s answers the same way:
hf download olaverse/PurpleMIST-Flash-1.0 serve.py purplemist.py --local-dir purplemist
python purplemist/serve.py --model olaverse/PurpleMIST-Flash-1.0 --port 8000
Flash runs on a single 24 GB GPU. PurpleMIST-Mini brings the same interface to 1.9B parameters. In English it
beats models up to four times its size, including Bongard-mini (7.5B) and Jeff-Gemma4-E2B (4.6B), with a median
latency of 46 ms. Both are open weights under Apache-2.0.
Benchmark your model
African Typed Decisions is open. Score your model on the all config, one request per case, and open a
discussion with your numbers. We’ll add you to the leaderboard. If you build for Hausa, Yorùbá, Swahili or any of
the other languages and want decision models that understand your users, we’d love to hear from you at
olaverse.co.uk.
Links: PurpleMIST-Flash-1.0 ·
PurpleMIST-Mini-1.0 ·
African Typed Decisions
About the benchmark: references come from LLM teachers, and the African cases are machine-written and
machine-checked. PurpleMIST was trained on separate data from the same pipeline. Full methodology is on the
dataset card.
Enjoyed this? Every model we write about is open weight and free to download.


