Ollama 支持基于 Jev API 的决策模型,新增 nimble 等三款模型
Ollama now supports Jev-style decision models
Ollama 0.35 通过新 /v1/systemone 端点支持基于 TypeSafe Jev API 的决策模型,可在本地一次请求回答多个命名问题,适合工单分诊、模型路由和内容审核等快速决策任务。
官方宣布本地运行决策模型,给出延迟数据和 API 用法,可帮助读者评估本地快速决策场景的可行性。
September 29, 2026
Ollama now supports decision models, based on TypeSafe’s Jev API for fast, typed decisions:
- No additional costs
- Lower latency when run locally
- Three new decision models available today via Ollama
This new API is available as of Ollama 0.35 by using the new /v1/systemone endpoint. Send text as state with a set of named questions, and a model running on your machine answers them all in one request. This is great for tasks that require fast decisions, such as ticket triage, model routing, and content or safety moderation.
Near-instant decisions
Decision models on Ollama are fast, as requests don’t have to travel over a network. Nimble 9B averaged 91ms per decision in the Pac-Man example below when running locally on an M5 Max. That’s fast enough to make rapid decisions such as playing a game or processing content in real time:
Move 1
Available models
Three new decision models are available to run via Ollama:
nimble: open-source 9B parameter decision model developed by Bespoke Labstev1: an experimental 4B decision model from Together AItev1:0.8b: an experimental 0.8B decision model from Together AI
More decision models are coming soon, including models served by Ollama’s cloud.
Decision models on Ollama
Bespoke Labs public benchmarks · accuracy, higher is better
Get started
To get started, first download or upgrade to the latest version of Ollama. Next, download a decision model such as nimble:
ollama pull nimble
You can make a request via curl or via TypeSafe’s official Python SDK.
Request
curl http://localhost:11434/v1/systemone -d '{
"model": "nimble",
"state": {
"ticket": "I was charged twice. Please refund the extra payment."
},
"questions": {
"team": {
"type": "choice",
"instructions": "Which team should handle this ticket?",
"criteria": {
"billing": "Payments and refunds",
"technical": "Bugs and integrations",
"other": "None of the above"
}
},
"refund": {
"type": "noul",
"instructions": "Does the customer explicitly ask for a refund?"
},
"urgency": {
"type": "score",
"instructions": "How urgent is this ticket?",
"criteria": ["Routine", "Soon", "Urgent"]
}
}
}'Response
{
"model": "nimble",
"answers": {
"team": {
"type": "choice",
"choice": "billing",
"probabilities": {"billing": 0.985, "technical": 0.012, "other": 0.003},
"confidence": 0.922
},
"refund": {"type": "noul", "noul": 0.997},
"urgency": {
"type": "score",
"score": 0.815,
"legend": {"0": "Routine", "1": "Soon", "2": "Urgent"},
"probabilities": {"0": 0.378, "1": 0.429, "2": 0.193},
"confidence": 0.046
}
},
"usage": {"input_tokens": 841, "output_tokens": 4}
}What’s next
This is the first of many releases to come adding decision model support to Ollama. Future updates will include:
- Faster performance on Apple Silicon powered by MLX
- More models specializing in different kinds of decision making
来源:Ollama:Blog · ollama.com