Ollama 0.35 adds native support for Jev-style decision models
TL;DR
Ollama 0.35 now supports Jev‑style decision models via a new /v1/systemone endpoint, enabling fast, typed decisions on‑device without additional fees.
New API endpoint for typed decisions
Conclusion: The /v1/systemone endpoint lets developers send a textual state and a set of named questions to a local model, receiving structured answers in a single request. This design follows TypeSafe’s Jev API and removes network round‑trips, lowering latency and cost.
POST http://localhost:11434/v1/systemone
{
"model": "nimble",
"state": "Our checkout has returned 500 errors since 9am.",
"questions": {
"label": {
"type": "choice",
"instructions": "Which label fits this ticket?",
"criteria": {"billing": null, "bug": null, "account": null}
}
}
}
The response returns the chosen answer, per‑choice probabilities, and a confidence score, together with token usage statistics.
Near‑instant local inference
Conclusion: Running decision models locally eliminates network latency; Ollama reports 91 ms per decision for the 9‑billion‑parameter nimble model on an Apple M5 Max, fast enough for real‑time game playing or content moderation.
The blog post demonstrates a Pac‑Man‑style move decision where the model chooses “left” with 0.65 probability in 91 ms.
Decision models currently available
Conclusion: Three decision‑model families are shipped with Ollama 0.35:
nimble– an open‑source 9 B‑parameter model from Bespoke Labs.tev1– an experimental 4 B‑parameter model from Together AI.tev1:0.8b– an experimental 0.8 B‑parameter model from Together AI.
Benchmarks published by Bespoke Labs show mean accuracy across 13 public data sets (3,880 decisions). The blog links to the full evaluation results and the benchmark suite for reproducibility.
Example workflow
Conclusion: Getting started requires only an upgrade to Ollama 0.35, pulling a decision model, and issuing a curl request (or using TypeSafe’s Python SDK). The example request classifies a support ticket, routes it to a team, detects a refund request, and scores urgency, returning structured JSON with choices, probabilities, and confidence values.
ollama pull nimble
curl http://localhost:11434/v1/systemone -d '{
"model": "nimble",
"state": {"ticket": "I was charged twice. Please refund the extra payment."},
"questions": {
"team": {"type": "choice", "instructions": "Which team should handle this ticket?", "criteria": {"billing": "Payments and refunds", "technical": "Bugs and integrations", "other": "None of the above"}},
"refund": {"type": "noul", "instructions": "Does the customer explicitly ask for a refund?"},
"urgency": {"type": "score", "instructions": "How urgent is this ticket?", "criteria": ["Routine", "Soon", "Urgent"]}
}
}'
The response includes a high‑confidence billing label (0.985 probability), a near‑certain refund flag (0.997), and an urgency score of 0.815 mapped to the “Soon” category.
Roadmap and future enhancements
Conclusion: Ollama plans to extend decision‑model support with faster Apple‑Silicon inference via MLX and to add more specialized models, both locally and in the cloud.
All performance numbers, model names, and benchmark references are taken directly from the Ollama blog post dated September 29 2026.