All News
liquid-aid1decision-modelsopen-weightsedge-aiclassificationroutingcalibrationnvidia-jetsonvision-language

Liquid AI's open d1: real decision models, smaller than the tease

Liquid AI opened d1-3B and d1-omni-600M, decision models that return calibrated probabilities in one pass. We separate the shipped artifact from the teaser.

Vlad MakarovVlad Makarovreviewed and published
5 min read
Liquid AI's open d1: real decision models, smaller than the tease

Liquid AI shipped two open-weight members of its d1 family on 7 October, and part of the local-model crowd that had spent the day expecting a general open release read the result as a letdown. The models, d1-3B and the experimental d1-omni-600M, are decision models: they do not write text. Given a state — text, JSON, images, or a mix — and a set of named questions, they return calibrated probabilities for each answer in a single forward pass, with zero output tokens. The artifact is real and narrow. The disappointment came from expecting something else.

What actually shipped

d1-3B is a 3.12B model built on LFM2.5-VL-3B, a decoder-only vision-language backbone. It reads text and images, runs to a 32,768-token context, and carries a SigLIP2 vision encoder. d1-omni-600M is the experimental sibling: 587M parameters on a bidirectional LFM2.5-Encoder-350M trunk, with a vision tower and a 17-layer FastConformer audio encoder attached. It accepts text with an image, or text with up to 30 seconds of speech, inside a 16,384-token context.

Both are on Hugging Face with day-one llama.cpp support, and both sit under the LFM Open License v1.0, a source-available weight license with commercial conditions rather than Apache or MIT. The model card is explicit that d1-3B "is not a chat model and does not write text." At check time its repository listed roughly 5,400 downloads and 186 likes.

What a decision model actually is

The mechanism is the interesting part. Unlike Liquid's generative models, these produce no tokens. The model reads the state and its questions once, and every answer is read directly off the distribution over the options. Liquid defines three question types: noul, a yes/no answered with a probability; choice, one label among named options with a probability per label; and score, a position on two to ten ordered levels. One request can ask several questions over the same state, so the input is read once for all of them.

That structure targets a specific job — routing and triage, moderation, intent and topic classification, extraction checks, reranking, LLM-as-a-judge scoring, agent guardrails, visual inspection. The pitch is economics: replace a chat-model call when the output you need is a structured decision, and generate nothing. Because there are no output tokens, Liquid measures end-to-end latency instead of tokens per second: 8 ms per question on an RTX 4090, 16 ms on a Jetson AGX Thor, 26 ms on a Jetson AGX Orin, 50 ms even on a Jetson Orin Nano. On an Apple M5 Pro it is 30 ms. Those are vendor numbers from warm calls, and they read as the most concrete claim in the release.

It is a second act, not a debut

Two days earlier, on 5 October, Liquid had introduced the hosted d1 model with text and image support through its API. That post is worth reading against this one: it frames d1 as a cost play against frontier models, reports quality and per-1,000-run costs across six applications, and closes by promising that Liquid plans to release open weights for upcoming models on Hugging Face soon. The 7 October release keeps that promise, but it adds downloadable members of the family rather than another way to reach the hosted service. Anyone who read the tease as the arrival of a general open Liquid model was reading a different announcement.

The scorecard is Liquid's own

Liquid reports d1-3B at 48.57 on the Decision Index 0.2.1 and calls it the best decision model under 10B. The claim is narrower than it sounds. On the same public leaderboard, Winnow-12B scores 50.02, so the top of the table is not d1-3B. Liquid scored its own entries with the official scorer and notes they are not a leaderboard submission; every other row is the leaderboard's.

ModelSizeDecision Index 0.2.1
Winnow-12B12B50.02
d1-3B3B48.57
Decider 35B-A3B36B47.11
JPT-9B9.7B46.89
d1-omni-600M587M15.95

Liquid also reports a mean of 82.9 for d1-3B across seven text tasks, but calls those internal evaluations based on public benchmarks, which leaves open the possibility that some test rows sat in training data. Liquid itself flagged the risk, dropping HelpSteer2 from d1-omni-600M over suspected overlap. There is no vision benchmark and no audio benchmark, because, as the model card puts it, audio decision benchmarks are currently an open problem. No independent evaluation of either model turned up. Even BenchLM, a third-party tracker that reproduces model results, carries only the six rows Liquid published and declines to give d1-3B a public rank, noting the model is not eligible.

The reaction was mostly "where is the use case"

The r/LocalLLaMA thread that opened the day, "New LFM to be released today," drew the skeptical response a vague tease invites. Reddit blocks third-party extraction, so the point stays qualitative: the visible replies wondered what a decision model is for, and Liquid's framing absorbed the hit. A same-day AI News digest summarized the mood by noting commenters were skeptical of the hype, with one saying "insane" has become synonymous with "mid." No engagement figure appears here because none could be verified.

The fair counter is that mid is the wrong word for the artifact. A 3B model that answers named questions with calibrated probabilities in single-digit milliseconds on a Jetson-class board is not something a buyer could get before. The complaint is about the category, not the quality: decision models look like a feature dressed as a product, and the crowd that shows up for open weights wants a general model it can chat with. The pattern repeats — the Strata inference-engine backlash three days earlier ran the same way, with a real tool buried under a borrowed adjective.

What would settle it

Two things are missing, and outsiders can check both. The first is an independent evaluation on the routing, classification and scoring tasks Liquid claims; DecisionBench and Fast Decisions are public datasets the models were measured against, so a third party can reproduce the numbers rather than trust them. The second is evidence that "decision model" is a market and not just a label. The comparison Liquid's own scorecard invites is telling: d1-3B beats a decision model twelve times its size, and neither reaches the top of the table. Aleph Alpha's Kolibri landed in a similar spot this month — an open artifact whose competitive claim rested on vendor numbers. If a general-purpose open weight was the expectation, this is not it. If a fast, calibrated classifier that fits on an edge board is useful, Liquid shipped exactly that, and the disappointment was mispriced.

Related Articles

Scroll down

to load the next article