Quick Answer:
On 29/09/2026 Liquid AI released d1, its first decision model. It classifies, routes and scores without generating a single output token, returning a probability for every possible answer instead of writing JSON. It offers three question types (Noul for yes/no, Choice for picking one option, Score for rating on an ordered scale), a 32,000-token context window, and access through Liquid's API and Vercel's AI Gateway. Secondary coverage says it is the first to top the Jev Decision Index on Hugging Face. Paid pricing, accuracy figures and latency figures have not been published, and the model is hosted-only, so treat it as a pilot candidate rather than a migration target.
A surprising amount of production AI is a very expensive way to say "yes", "no" or "billing". A team asks a large language model whether a message is spam, tells it to answer in JSON, pays for every token of that JSON, then writes code to parse it and handle the days it comes back malformed. Liquid AI's d1 is built on the observation that, for that kind of question, the writing step is pure overhead.
This review explains what d1 actually is, how its three question types work, what the developer documentation and independent coverage confirm, which claims remain unverified, and how a sensible team should test it. Where Liquid or its reporters have not published a number, we say so rather than guess.
Sources checked for this article include DataNorth's Liquid AI releases d1 report, AlphaSignal's d1 write-up, the LLM Reference d1 listing, and AI Weekly's coverage of the Jev Decision Index. Liquid AI's own developer documentation is the origin of most facts, but we could not retrieve a dedicated launch blog post, because DataNorth reports there is none. Figures below are attributed to the outlet that reported them.
Julian Goldie SEO's quick take on the d1 release, worth watching alongside the documentation-based detail in this review.
Executive Summary
- Released 29/09/2026 (per DataNorth and AlphaSignal): Liquid AI's first decision model, a specialist alternative to a language model for classification, routing, approval gates and scoring.
- Zero generated output tokens: d1 returns typed answers with probabilities and reports
usage.output_tokensas 0, so billing is for input only. - Three question types: Noul (yes/no, a probability from 0 to 1), Choice (pick one, a probability per option) and Score (rate on an ordered scale, a weighted position).
- 32,000-token context window, reachable through Liquid's API (model name
d1:free) and Vercel's AI Gateway (liquid/d1). OpenRouter support is planned but undated. - Reported benchmark lead: AlphaSignal and others say d1 is number one on the Jev Decision Index, a Hugging Face-hosted leaderboard covering 132,422 requests across 37 benchmarks. The previous leader, Jev 1.13, scored a composite near 74.4. d1's own score was not in the sources we could read.
- Targeted use cases: content moderation, ticket triage, guardrails and approving agent tool calls.
- Unpublished: paid pricing, accuracy and latency figures, parameter count, and a launch blog post. DataNorth's verdict is "worth testing, not yet worth committing to".
- Hosted only: no downloadable weights, so it is unsuitable for data that cannot leave your network.
The Problem d1 Solves
Most teams that use a language model for classification follow the same pattern. They write a prompt describing the labels, ask the model to reply with a small piece of JSON, and then parse the reply into whatever their software expects. DataNorth describes the cost of this plainly: the model produces that text one token at a time, and you pay for every one of them. d1 skips the writing step entirely. It reads your input and returns numbers.
AlphaSignal frames the same idea in engineering terms. Language models generate output token by token even when the application needs only a boolean or a label. By skipping that decoding stage, d1 aims to give a more predictable latency profile and to remove a family of failure modes from the path entirely: malformed JSON, missing fields and invalid enum values. In the usual approach those problems are handled with retries, validators and defensive parsing code. With a typed response they cannot arise in the same way, because the service returns structured data without producing it through autoregressive text generation.
There is a second, subtler benefit. A language model asked a yes/no question usually returns a single word and hides its uncertainty. d1 returns a probability, which means your software can set a threshold. A moderation system might auto-remove content above 0.95, queue content between 0.60 and 0.95 for human review, and let everything else through. That is a far more useful interface for production systems than a bare label, provided the probabilities are well calibrated on your data, which is the part you must test for yourself.
None of this makes d1 a general chatbot. It cannot draft an email, summarise a report or write code. It is a narrow tool aimed at one slice of AI workloads, and the case for it rests on that slice being large. For many businesses it is: classification, routing and approval logic sit behind a high proportion of automated workflows, from inbox handling to fraud queues.
Lineage: Liquid AI and Its Specialist Models
Liquid AI is a Boston-based company spun out of MIT's Computer Science and Artificial Intelligence Laboratory in 2023. It is best known for its Liquid Foundation Models, small, efficient models designed to run on phones, laptops and embedded hardware. Our earlier review of LFM2.5-DSpark covers the company's background and its speculative-decoding accelerators in detail.
d1 is a different kind of release. LFM models are open-weight language models you can download. d1 is a hosted decision model with no published weights. LLM Reference lists it as proprietary hosted weights, with Hugging Face, GGUF, MLX and ONNX formats unavailable and the model not trainable by users. That is a notable departure for a company whose reputation was built on small models you could run yourself, and it is the single most important thing to understand before planning around it.
Timing matters too. DataNorth observes that Liquid's own news page still ended at LFM2.5-VL-DSpark on 24/09/2026, and that Liquid wrote no launch post for d1. Everything about it sits in developer documentation, which describes the API and, in DataNorth's words, gives no accuracy figures, no latency figures and no price. One listing, LLM Reference, shows a release date of 22/09/2026, while DataNorth and AlphaSignal's coverage point to 29/09/2026. We use 29/09/2026 because two outlets agree on it, but the discrepancy is a reminder that this was a quiet documentation-first release rather than a polished launch.
How d1 Works: Noul, Choice and Score
d1 exposes three primitives, which are the only shapes of answer it will give. The restriction is the point: because the answer type is fixed in advance, the service can return numbers directly instead of generating text and hoping it parses. You can include several questions in a single request, each with its own type, and the response contains an answer for each.
Noul: yes or no, with a probability
A Noul is a yes/no question answered with a probability between 0 and 1. DataNorth's example is "Is this message spam?" returning 0.92. AlphaSignal's Python example asks whether a customer message is a complaint and prints a value of 0.999. The name is Liquid's own coinage, and the sources we read do not explain its origin, so we will not speculate. Practically, Noul is the building block for any boolean gate: is this toxic, is this a refund request, is this tool call safe to run.
Choice: routing between named options
A Choice is a pick from named options, answered with a probability for each. The documentation example is a support ticket routed to billing, technical or account, which might return 0.65, 0.30 and 0.05. Notice what this gives you that a single label does not. A ticket at 0.65 billing and 0.30 technical is genuinely ambiguous, and your workflow can respond by sending it to a human or to a generalist queue instead of confidently misrouting it.
Score: a position on an ordered scale
A Score is a rating on an ordered scale, answered with a weighted position. The example is an urgency rating that comes back as 1.85, sitting between "medium" and "high". Because the result is a position rather than a bucket, you can sort a queue by it, plot it over time or apply different thresholds for different customers. It suits ordinal judgements such as urgency, sentiment intensity, relevance or quality.
The API in Practice
According to AlphaSignal, the endpoint is https://api.liquid.ai/decisions/v1/systemone, and calls use the model name d1:free. A Liquid API key is required, obtained through the API console. The published Python example uses a TypeSafe SDK client and a Noul question named is_complaint; the result is read from result.answers["is_complaint"].noul. The response includes input-token usage and reports usage.output_tokens as 0.
Two practical points follow. First, because there are no output tokens, usage is driven primarily by the size of the input and the question definitions. Long documents with many questions cost more than short messages with one. Second, DataNorth quotes Liquid's framing that d1 "returns a probability for every possible answer, so you are billed for input only". For a workload where a conventional model would write 30 to 100 tokens of JSON per call, removing that output is the entire economic argument. Whether it is a good argument depends on the unpublished paid rate, which we return to below.
d1 is also reachable through Vercel's AI Gateway under the model ID liquid/d1, which matters for teams that already route model traffic through a gateway for logging, fallbacks and spend controls. A gateway route means you can trial d1 by changing a model string in an existing integration rather than adding a new vendor SDK. The 32,000-token context window is enough for most messages, tickets, documents and conversation histories, but it is far smaller than the contexts of frontier chat models, so very long inputs would need truncating or chunking.
Use Cases: Moderation, Triage, Guardrails and Tool-Call Approval
The sources name four target workloads. They share a structure: a fixed schema, high volume, and an answer that a program rather than a person will consume.
- Content moderation. A Noul per policy category (harassment, spam, self-harm, and so on) gives you a probability for each, which you can threshold separately. Policies differ in tolerance for false positives, and a probability lets you encode that.
- Ticket triage. A Choice for the queue and a Score for urgency can run in one request, giving a routing decision and a priority ordering from a single call.
- Guardrails. A Noul asking whether an input is a prompt-injection attempt, or whether an output leaks personal data, can sit in front of or behind a larger model. AlphaSignal reports that d1 improves on prior models in prompt-injection robustness, which is the relevant property here, though no figure was published.
- Approving agent tool calls. As agents gain the ability to run commands, send messages and spend money, someone has to decide whether each proposed action is allowed. A fast Noul such as "is this action consistent with the user's request and policy?" makes a natural approval gate. Our coverage of always-on agents such as OpenAI Dots, Grok Bot and Hermes Agent shows how quickly autonomous actions are multiplying, and every one of them needs an approval layer that is cheap enough to run on every step.
The agent case is the most interesting strategically. An approval gate that costs as much as the agent's main model call tends to get switched off in production. One that adds only input-token cost and a predictable latency is far more likely to stay on. It is also where calibration matters most, since a gate that is confidently wrong in the permissive direction is worse than no gate. For coding agents specifically, compare how harnesses such as the DeepSeek Harness and tools like Claude Code handle permission prompts today, usually by asking a human.
The Jev Decision Index Claim
The headline claim is that d1 is the first to beat "Jev" and take number one on the Jev Decision Index. Here is what is confirmed about the index itself, from AlphaSignal and AI Weekly:
- It is a leaderboard posted to Hugging Face by multimodalart, scoring decision models against a frozen suite of 132,422 requests across 37 benchmarks.
- Nineteen of those benchmarks form the scored panel, split equally across five areas: Tools & Automation, Retrieval & Classification, Language Understanding, Knowledge & Reasoning, and Arts & Human Judgment.
- The headline index is 100 times the mean of the five area scores, giving a value from 0 to 100.
- The rules are strict: engines cannot truncate requests, cannot drop options, and get one fixed prompt rendering across the suite. Failed requests score zero before normalisation.
- Before d1, Jev 1.13 led with a composite near 74.4, according to AlphaSignal's reading of the leaderboard.
- AI Weekly reported the index as ranking 30+ open-weight decision models, and noted it welcomes submissions by pull request with dataset links and hardware details.
What is not confirmed is d1's own composite score, its margin over Jev, or its per-area results. None of the sources we could read quoted them, so we do not either. There is also a methodological question. DataNorth notes that the index ranks open-weight models and d1 has no weights, and says that until Liquid or the leaderboard's maintainer explains how an API-only model was measured, the placement "is not evidence you should act on". That is a fair caution. It does not mean the claim is false, only that the evidence is thin and secondhand.
AlphaSignal adds the more general caveat that a frozen request set makes comparisons consistent, but a leaderboard cannot predict production accuracy, calibration, latency or cost for your specific label set and traffic. We agree. A model can top a 37-benchmark average and still do worse than a cheaper one on your three labels. The benchmark is a reason to put d1 on your shortlist, not a reason to skip your own evaluation.
AlphaSignal also lists the improvements Liquid claims over earlier models: stronger multilingual evaluation results, greater prompt-injection robustness, better handling of longer inputs, and faster structured decisions. Again, these are directional claims with no published numbers.
The Road Decider Demo
Liquid's cookbook includes a demonstration called Road Decider, a pixel-art driving game in which a model acts as the controller. For every frame, the model selects LEFT, CENTER or RIGHT and returns a confidence score. The idea is to show that d1 can make a fresh decision on each frame in real time, something that would be impractical with a model that has to write out a reply each time.

The screenshot shows the two models side by side. At the moment captured, d1 has all three lives and reports a confidence of 1.00 on CENTER, while Jev has two lives and reports 0.92 on RIGHT. It is tempting to read that as a win, but it is one frame from a game, and Liquid has not published aggregate results from it. Its value is as an illustration of the response shape: a discrete choice plus a confidence value, delivered quickly enough to drive a loop.
Access, Pricing and What Is Not Published
This is the section where honesty matters most, because the gaps are the story. Here is a plain accounting.
- Published: a 32,000-token context window; access through Liquid's API and Vercel's AI Gateway (
liquid/d1); a free tier under the named1:free; billing on input only; OpenRouter support planned. - Not published: the paid rate. DataNorth states that Liquid has not said what d1 costs once you leave the free tier, so a business case cannot be calculated. LLM Reference lists input and output prices as free, reflecting the free tier only.
- Not published: accuracy figures, latency figures, calibration data, parameter count, rate limits as a general specification, and data-handling terms. AlphaSignal advises production users to confirm current rate limits, data-handling terms, regional availability and paid pricing in the console before deployment.
- Not available: model weights, fine-tuning, and any self-hosted option.
Because we cannot give a sourced price, we will not convert to pounds or compare it with GPT-class or Claude-class API rates. Anyone telling you precisely how much d1 saves is guessing. What can be said logically is that removing output tokens helps most where output is the larger part of the bill, and helps least where inputs are long. A call that sends a 10,000-token document and asks one question will see a smaller relative saving than a call that sends a 50-word message and asks five.
Limitations
- No paid pricing. You cannot forecast cost beyond the free tier.
- No accuracy, latency or calibration data from Liquid. The case rests on a reported leaderboard position.
- Leaderboard methodology is unclear for an API-only model, as DataNorth points out.
- Hosted only. Data leaves your network, which rules d1 out for many regulated or sensitive workloads.
- Narrow by design. It answers Noul, Choice and Score questions only. Anything needing explanation, free text or reasoning shown to a user needs a different model.
- No explanations. A probability tells you how confident the model is, not why. If you need an audit trail of reasons for a moderation decision, you will have to add a separate process.
- 32,000-token ceiling. Adequate for most classification inputs, restrictive for whole-codebase or long-document work.
- Quiet launch. No launch post, a conflicting release date on one listing, and OpenRouter availability undated. Documentation-only releases can change without announcement.
- Calibration is not guaranteed. The pitch of "calibrated probabilities" is Liquid's claim. Whether 0.9 really means right nine times in ten on your data is something only a test can show.
How It Compares
Versus a general language model returning JSON. This is the comparison d1 is designed to win. A frontier model can do the same classification and also explain itself, but it costs output tokens, adds latency and can return invalid structure. d1 trades flexibility for a fixed interface and input-only billing. Whether the trade pays off depends on volume and on that unpublished price.
Versus Fastino's GLiNER2.5-Decide. DataNorth says Fastino released this on 25/09/2026, four days earlier, aimed at the same work. The meaningful difference is not the output format but where the model runs. Fastino ships weights you can host yourself; Liquid keeps d1 behind its API. For a team already sending this traffic to a hosted model, DataNorth says that difference costs nothing and d1 is a straight swap. For a team classifying regulated or customer data inside its own network, it rules d1 out completely, and no price cut will change that.
Versus Liquid's own LFM models. Small LFM models can be run locally and fine-tuned, and our LFM2.5-DSpark review shows how far Liquid is pushing local speed. A fine-tuned small model remains a credible alternative for teams that want ownership of the classifier. d1 is the opposite bet: a managed service where Liquid has done the specialisation for you.
Versus rule-based or classical classifiers. A traditional trained classifier can be faster and cheaper still for a single stable task, but needs labelled data and maintenance. d1 sits between that and a full language model: no training needed, flexible labels defined per request, with probabilities. For rapidly changing label sets, that flexibility is the draw.
Who Should Use It
Worth piloting now: teams that already send high-volume classification, routing or moderation traffic to a hosted language model and generate JSON just to get a label. Developers building agent approval layers who want a cheap, structured gate. Anyone with a gateway in place who can trial liquid/d1 by changing a model string.
Wait or look elsewhere: organisations whose data cannot leave their own servers; teams that need a price before they can get budget approval; anyone who needs written reasoning alongside each decision; and workloads dominated by very long inputs, where input-token cost dominates and the output saving is small.
A One-Week Pilot Plan
DataNorth's advice is to take the single classification route with the highest volume, run it against d1 and your current model for a week, and compare accuracy and cost side by side, which it estimates at about a day of engineering work. We would add a few specifics, drawing on AlphaSignal's guidance to test the full decision pipeline rather than compare model scores alone.
- Pick one route with clear labels and a source of truth, such as tickets that humans later re-categorised.
- Shadow-run d1 alongside the current model on live traffic without acting on its output.
- Measure accuracy per class, not just overall, and check how often probabilities above your threshold are actually right.
- Check calibration: bucket predictions by probability and compare to observed correctness.
- Record latency and token usage, including input-token counts, so you can model cost once pricing is published.
- Set thresholds and a human-review band using the probabilities, and test what happens on borderline cases.
- Monitor drift after launch: class frequency, probability distributions, error rates, latency and input-token usage, as AlphaSignal recommends.
- Test adversarially if it will act as a guardrail, using prompt-injection samples relevant to your product.
The Bottom Line
d1 is a sensible idea executed as a quiet, documentation-only release. The core proposition is strong: if all you need is a label, a boolean or a score, you should not be paying a language model to write JSON and then parsing it. A typed answer with a probability is a better interface for software, and removing output tokens should help on cost and latency where the workload fits.
But the evidence is thin. There is no launch post, no published price beyond a free tier, no accuracy or latency numbers, and a leaderboard claim that rests on secondary reporting and an unexplained methodology for an API-only model. The hosted-only design also closes the door on sensitive data. The right response is the one DataNorth reached: pilot it on your busiest classification route, measure it against your current setup, and do not plan anything wider until Liquid publishes the numbers. Worth testing, not yet worth committing to.
Sources
- DataNorth: Liquid AI releases d1 (release date, access, context window, pricing gaps, Fastino comparison; lead image credit).
- AlphaSignal: Liquid AI's d1 Makes Decisions Without Generating a Single Token (API endpoint, Jev Decision Index detail, Road Decider demo, pilot guidance; demo screenshot).
- LLM Reference: d1 on Liquid AI (hosted-only status, free listing, structured outputs).
- AI Weekly: Jev Decision Index ranks 30+ open-weight decision models (index design and rules).
- Liquid AI (company site and developer documentation).
- Vercel AI Gateway (gateway through which d1 is offered as liquid/d1).
Last updated: 04/10/2026. Sourced from DataNorth, AlphaSignal, LLM Reference and AI Weekly. d1's paid pricing, accuracy and latency figures were not published at the time of writing, and this article will be revised when Liquid AI releases them.
Get the free guide: Claude vs ChatGPT, Gemini & Grok
A 20-page playbook covering everything you need to choose and use the big four AI models in 2026, full cost and feature comparisons, what each is best (and worst) at, and how-tos for images, vectors, building a website, Claude Code and more.








