Skip to main content
Instant answers, edge-ready

Multimodal Open d1 Decision Models for the Edge

Liquid AI has unveiled two open-weight decision models designed to operate efficiently at the edge. Unlike conventional generative models that produce tokens sequentially, these models deliver answers in a single forward pass—a fundamental shift that unlocks deployment possibilities on resource-constrained devices. The approach promises speed and efficiency where it matters most.
Multimodal Open d1 Decision Models for the Edge
Multimodal Open d1 Decision Models for the Edge

Liquid AI has released two open-weight decision models: d1-3B and d1-omni-600M (experimental). Unlike generative models that produce tokens sequentially, decision models answer questions in a single forward pass, making them suitable for edge deployment.

d1-3B is trained from LFM2.5-VL-3B, a decoder-only vision language model that accepts text and images. d1-omni-600M is trained from LFM2.5-Encoder-350M, a bidirectional encoder that adds vision and audio encoders, accepting either text and images or text and audio as inputs. The d1-omni-600M model is currently in early research release and undergoing further development.

Benchmark Performance

— Hugging Face

On the Decision Index 0.2.1, d1-3B scores 48.57, outperforming every 4B and 9B model and Decider 35B-A3B (47.11). Across seven public datasets spanning reading comprehension, toxicity detection, intent classification, medical QA, and cross-lingual understanding, d1-3B achieves a mean score of 82.9—the highest in its class and above Decider 4B (81.1). d1-omni-600M scores 78.4, surpassing Decider 2B (77.1) with only a quarter of the parameters.

Benchmark d1-omni-600M d1-3B Decider 2B Decider 4B
SQuAD 2.0 74.0 83.3 67.7 76.0
Civil Comments 95.8 93.3 93.6 92.8
MASSIVE intent 86.1 86.9 81.1 88.3
PubMedQA 61.3 68.3 65.7 63.3
BoolQ 77.7 86.3 87.3 89.0
XNLI 74.7 85.6 85.0 88.6
PAWS-X 79.5 76.4 59.5 69.8
Mean 78.4 82.9 77.1 81.1

The models retain their backbone capabilities: d1-3B preserves the vision capabilities of LFM2.5-VL-3B on standard vision benchmarks, and d1-omni-600M handles all three modalities. Vision and audio benchmarks are not reported, as the Decision Index v0.3 includes only a private vision split and audio decision benchmarks remain an open problem.

Edge Inference Performance

In collaboration with NVIDIA, d1-3B was evaluated across NVIDIA GeForce RTX 4090, NVIDIA Jetson AGX Thor, Jetson AGX Orin 64 GB, and Jetson Orin Nano. d1-omni-600M speed numbers are not reported in this release. d1-3B answers a single question in under 50 ms on every measured device. Processing three questions takes only 1.3x the time of one, with the AGX Thor going from 16 ms to 20 ms.

One question 3 questions 3.4K-token state 384px image 64 states, packed
Apple M5 Pro 30 ms 41 ms 640 ms 62 ms 78 / s
Jetson AGX Thor 16 ms 20 ms 220 ms 35 ms 262 / s
Jetson AGX Orin 64 GB 26 ms 35 ms 560 ms 83 ms 110 / s
Jetson Orin Nano 50 ms 73 ms 1,640 ms 202 ms 38 / s

On GPU, d1-3B answers a question in under 10 ms and processes a 384px image in under 18 ms on both platforms tested.

One question 3 questions 3.4K-token state 384px image 64 states, packed
NVIDIA RTX 4090 8 ms 21 ms 102 ms 17 ms 475 / s
AMD MI325X 9 ms 14 ms 44 ms 18 ms 1,106 / s

Getting Started

Both models are open-weight and available on Hugging Face. To use d1-3B, install dependencies (requires transformers>=5.14):

pip install "transformers>=5.14" torch torchvision pillow

Load the model with trust_remote_code=True:

import io import urllib.request import torch from PIL import Image from transformers import AutoModel device = "cuda" if torch.cuda.is_available() else "mps" if torch.backends.mps.is_available() else "cpu" model = AutoModel.from_pretrained("LiquidAI/d1-3B", trust_remote_code=True, dtype=torch.float32 if device == "cpu" else torch.bfloat16).to(device) questions = { "refund": {"type": "noul", "instructions": "Is the customer asking for a refund?"}, "team": {"type": "choice", "instructions": "Which team should handle this?", "criteria": {"billing": "Charges, refunds, invoices", "technical": "App or site faults", "fraud": "Suspected unauthorised use"}}, "urgency": {"type": "score", "instructions": "How urgent is this?", "criteria": ["Can wait", "Today", "Blocking the customer now"]}, } print(model.system_one("I was charged twice this month, please refund one of them.", questions)) url = "http://images.cocodataset.org/val2017/000000039769.jpg" photo = Image.open(io.BytesIO(urllib.request.urlopen(url).read())) print(model.system_one(None, {"cats": {"type": "choice", "instructions": "How many cats are there?", "criteria": {"one": "One", "two": "Two", "more": "Three or more"}}}, images=[photo])) tickets = ["Where is my parcel? It was due Monday.", "The app crashes when I open settings."] print(model.system_one_batch([(t, {"team": questions["team"]}) for t in tickets]))

For d1-omni-600M instructions, see the model card. Both models are available for download on Hugging Face, and demos can be tried in the System One Arcade Hugging Face Space.

Most Read

What the community is saying
community speaking about this topic
@mwaldman130 unknown likes · 2026-10-09
3 weeks later, @liquidai shipped a 3B that goes head to head with Jev -- on a Mac, 30 ms latency, and *with* vision...Today we release Open d1: two open-weight multimodal models in our d1 decision mod
@fahdmirza unknown likes · 2026-10-09
d1-omni-600M: a 587M model that decides from text, images AND speech. One small model, one forward pass, no text generation. Vision test: spotted two phones and distress in the image at 99.9%.

Felipe Santos

“Artificial intelligence can process the world in milliseconds, but only the human heart can give meaning to every second lived” – Mr. Santos