Ollama 0.35 新增對 Jev 風格決策模型的原生支援
TL;DR
Ollama 0.35 現在透過新的 /v1/systemone 端點支援 Jev 風格的決策模型,可在裝置上快速、類型安全地進行決策,且無額外費用。
新的 API 端點用於類型化決策
結論: /v1/systemone 端點讓開發者可以將文字形式的 狀態 和一組命名問題送至本地模型,並在單一請求中獲得結構化回應。此設計遵循 TypeSafe 的 Jev API,消除網路往返,降低延遲與成本。
POST http://localhost:11434/v1/systemone
{
"model": "nimble",
"state": "Our checkout has returned 500 errors since 9am.",
"questions": {
"label": {
"type": "choice",
"instructions": "Which label fits this ticket?",
"criteria": {"billing": null, "bug": null, "account": null}
}
}
}
回應會返回所選答案、每項選擇的機率以及信心分數,同時附帶 token 使用統計資料。
近乎即時的本地推理
結論: 在本地運行決策模型可消除網路延遲;Ollama 報告指出,在 Apple M5 Max 上,90 億參數的 nimble 模型每項決策僅需 91 ms,快到足以實時應用於遊戲或內容審核。
部落格文章示範了一個類 Pac-Man 的移動決策,模型在 91 ms 內以 0.65 的機率選擇「左」。
目前可用的決策模型
結論: Ollama 0.35 隨附三種決策模型家族:
nimble– 由 Bespoke Labs 提供的開源 9 B 參數模型。tev1– 由 Together AI 提供的實驗性 4 B 參數模型。tev1:0.8b– 由 Together AI 提供的實驗性 0.8 B 參數模型。
Bespoke Labs 發布的基準測試顯示,這三種模型在 13 個公開資料集(共 3,880 次決策)上的平均準確率。部落格連結至完整的評估結果與可重現的基準套件。
實例工作流程
結論: 開始使用只需升級至 Ollama 0.35,拉取一個決策模型,並發出 curl 請求(或使用 TypeSafe 的 Python SDK)。範例請求可分類支援工單、路由至團隊、偵測退款要求,並評估緊急程度,回傳包含選擇、機率與信心值的結構化 JSON。
ollama pull nimble
curl http://localhost:11434/v1/systemone -d '{
"model": "nimble",
"state": {"ticket": "I was charged twice. Please refund the extra payment."},
"questions": {
"team": {"type": "choice", "instructions": "Which team should handle this ticket?", "criteria": {"billing": "Payments and refunds", "technical": "Bugs and integrations", "other": "None of the above"}},
"refund": {"type": "noul", "instructions": "Does the customer explicitly ask for a refund?"},
"urgency": {"type": "score", "instructions": "How urgent is this ticket?", "criteria": ["Routine", "Soon", "Urgent"]}
}
}'
回應包含高信心度的帳務標籤(0.985 機率)、近乎確定的退款旗標(0.997),以及映射至「 Soon」類別的緊急程度分數 0.815。
路線圖與未來增強
結論: Ollama 計畫透過 MLX 提升 Apple Silicon 的推理速度,並增加更多專用模型,無論是本地還是雲端皆可支援。
- 所有效能數字、模型名稱與基準參考均直接取自 2026 年 9 月 29 日發布的 Ollama 部落格文章。