ollama

Skip to main content

Tag: ollama

Simple black-and-white line drawing of a llama's face with large eyes and upright ears.

Ollama now supports Jev-style decision models

Ollama now supports decision models based on TypeSafe's Jev API for fast, typed decisions. The new API is available as of Ollama 0.35 using the /v1/systemone endpoint. Send text as state with a set of named questions, and a model running locally answers them all in one request. This is useful for tasks requiring fast decisions, such as ticket triage, model routing, and content or safety moderation. curl http://localhost:11434/v1/systemone -d '{ "model": "nimble", "state": "Our checkout has returned 500 errors since 9am.", "questions": { "label": { "type": "choice", "instructions": "Which...

Continue reading

Google Cloud logo above text "Startups open up to Gemma" beside a flowing, translucent abstract sculpture in warm hues.

Building Smarter AI Stacks: Why Startups Are Ditching One-Size-Fits-All Models

Every week, I talk with founders building at an unbelievable pace. Teams are moving from inception to product-market fit faster than ever, with foundation models wired deeply into their core product workflows. Yet as startup architectures mature, a clear divide has emerged between teams struggling with margins and those scaling sustainably. The most effective engineering teams have abandoned the one-size-fits-all model strategy. In the early days of LLMs, the default architecture was simple: send every interaction to the largest model available. But as applications move into production,...

Continue reading

Frontier Intelligence Goes Local

Frontier Intelligence Goes Local

At IFA 2026, NVIDIA, Microsoft and their partners are unveiling faster inference capabilities and simplified tools for running AI agents locally on NVIDIA hardware. New compact NVIDIA RTX Spark Windows PCs arrive in October, offering AI enthusiasts, developers and creators local, secure agent deployment. Simplified Local AI Support and Performance Improvements Simplified local AI support for NVIDIA GPUs is coming to Hermes Agent, OpenClaw and Perplexity Portable Computer. Up to 1.9x faster local inference — New llama.cpp and vLLM optimizations are available now directly through LM Studio...

Continue reading

Glowing green network links connect laptops and devices in a distributed system on a dark reflective surface.

NVIDIA PAIR: Distributed AI Inference for Home Networks

AI agents are learning to work together more effectively. A lead agent can break complex tasks into smaller jobs and assign them to specialized subagents, with users increasingly running multiple agent sessions simultaneously. Multi-agent workflows are becoming more common, but this approach can bottleneck the system as many requests flood the GPU at once. NVIDIA Personal AI Router (PAIR) addresses this problem by leveraging local hardware to distribute inference requests across available systems on a home network. PAIR routes each independent request to an available machine and works with...

Continue reading