Skip to main content
Milliseconds, locally

Ollama now supports Jev-style decision models

Ollama now lets developers run decision models locally at blazing speeds. The new `/v1/systemone` endpoint supports Jev-style APIs, answering multiple typed questions in a single request—perfect for ticket triage, content moderation, and real-time tasks that demand instant answers.
Simple black-and-white line drawing of a llama's face with large eyes and upright ears.
Simple black-and-white line drawing of a llama's face with large eyes and upright ears.

Ollama now supports decision models based on TypeSafe's Jev API for fast, typed decisions. The new API is available as of Ollama 0.35 using the /v1/systemone endpoint. Send text as state with a set of named questions, and a model running locally answers them all in one request. This is useful for tasks requiring fast decisions, such as ticket triage, model routing, and content or safety moderation.

curl http://localhost:11434/v1/systemone -d '{ "model": "nimble", "state": "Our checkout has returned 500 errors since 9am.", "questions": { "label": { "type": "choice", "instructions": "Which label fits this ticket?", "criteria": {"billing": null, "bug": null, "account": null} } } }'

Near-instant decisions

Decision models on Ollama are fast since requests don't travel over a network. Nimble 9B averaged 91ms per decision in a Pac-Man example when running locally on an M5 Max—fast enough for rapid decisions such as playing games or processing content in real time:

{ "answers": { "move": { "type": "choice", "choice": "left", "probabilities": { "left": 0.65, "right": 0.35 }, "confidence": 0.07 } } }

Available models

Three new decision models are available via Ollama:

  • nimble: an open-source 9B parameter decision model from Bespoke Labs
  • tev1: an experimental 4B decision model from Together AI
  • tev1:0.8b: an experimental 0.8B decision model from Together AI

More decision models are coming soon, including models served by Ollama's cloud.

Mean accuracy across 13 public datasets with human labels covering 3,880 decisions shows competitive performance. Nimble and Tev1 were evaluated on Ollama; Jev 1.13 data is from Bespoke Labs' published run on the same decisions.

Get started

First, download or upgrade to the latest version of Ollama. Next, download a decision model such as nimble:

ollama pull nimble

You can make requests via curl or TypeSafe's official Python SDK.

Request

curl http://localhost:11434/v1/systemone -d '{ "model": "nimble", "state": { "ticket": "I was charged twice. Please refund the extra payment." }, "questions": { "team": { "type": "choice", "instructions": "Which team should handle this ticket?", "criteria": { "billing": "Payments and refunds", "technical": "Bugs and integrations", "other": "None of the above" } }, "refund": { "type": "noul", "instructions": "Does the customer explicitly ask for a refund?" }, "urgency": { "type": "score", "instructions": "How urgent is this ticket?", "criteria": ["Routine", "Soon", "Urgent"] } } }'

Response

{ "model": "nimble", "answers": { "team": { "type": "choice", "choice": "billing", "probabilities": {"billing": 0.985, "technical": 0.012, "other": 0.003}, "confidence": 0.922 }, "refund": {"type": "noul", "noul": 0.997}, "urgency": { "type": "score", "score": 0.815, "legend": {"0": "Routine", "1": "Soon", "2": "Urgent"}, "probabilities": {"0": 0.378, "1": 0.429, "2": 0.193}, "confidence": 0.046 } }, "usage": {"input_tokens": 841, "output_tokens": 4} }

Setup

uv add typesafe-sdk export TYPESAFE_BASE_URL=http://localhost:11434 export TYPESAFE_API_KEY=ollama export TYPESAFE_DEFAULT_MODEL=nimble

Request

from typesafe_sdk import Choice, Noul, Score, TypeSafeClient ticket = "I was charged twice. Please refund the extra payment." questions = { "team": Choice( instructions="Which team should handle this ticket?", criteria={ "billing": "Payments and refunds", "technical": "Bugs and integrations", "other": "None of the above", }, ), "refund": Noul( instructions="Does the customer explicitly ask for a refund?", ), "urgency": Score( instructions="How urgent is this ticket?", criteria=["Routine", "Soon", "Urgent"], ), } with TypeSafeClient(timeout=120) as client: result = client.system_one( state={"ticket": ticket}, questions=questions, ) print(result.choices["team"].choice) print(result.nouls["refund"].noul) print(result.scores["urgency"].score)

What's next

Future updates will include faster performance on Apple Silicon powered by MLX and more models specializing in different kinds of decision making.

Ollama’s Jev-style decision models are available to run locally at no additional cost. The launch lists Nimble, a 9B open-source model from Bespoke Labs, plus Together AI’s experimental Tev1 models in 4B and 0.8B sizes; they are distributed through Ollama’s model library rather than sold as physical products.

Model Price Where to buy As-of date
Nimble 9B No additional cost Ollama model library Sep. 2026
Tev1 4B No additional cost Ollama model library Sep. 2026
Tev1 0.8B No additional cost Ollama model library Sep. 2026
Prices subject to change.

Felipe Santos

“Artificial intelligence can process the world in milliseconds, but only the human heart can give meaning to every second lived” – Mr. Santos