AI Tools Review
Microsoft-Decision-1: The 9B Decision Model Explained

Insights

Microsoft-Decision-1: The 9B Decision Model Explained

AI Tools Review Editorial Team11 October 2026

    Quick Answer:

    Microsoft-Decision-1 is a small 9-billion-parameter decision model (a post-trained Qwen3.5-9B) that Microsoft launched in Foundry on 09/10/2026. Instead of writing text, it takes a fixed set of options and returns a calibrated probability for each in one pass. Microsoft reports the highest average accuracy across 36 benchmarks (83.5%), a median latency of 85 ms against 3.01 s for GPT-6 Sol (about 35x faster), and a price of $0.042 per million input tokens with free output. Almost every number is Microsoft's own, so treat it as a strong, unverified claim.

    The most interesting AI model of the week is not another frontier chatbot. It is a small model that refuses to talk. Microsoft-Decision-1 does one thing: you hand it a question and a list of allowed answers, and it tells you how likely each answer is, in about a tenth of a second.

    That sounds modest until you count how many decisions an AI agent makes per task. Every route, retry, approval, label and safety check is a decision, and today most of them are made by a large language model that generates hundreds of tokens to reach a one-word answer. Microsoft is betting that this layer deserves its own cheap, fast, calibrated model.

    AI Revolution X covers Microsoft Decision-1 from the 7:39 chapter, after the Gemini 4 Carbon leak.

    Executive Summary

    What it is: Microsoft-Decision-1 is a "decision model". Microsoft describes decision models as purpose-built to deliver structured outputs that software can act on immediately. You supply a prompt and a closed set of options; the model returns a probability for each option through a structured API call. It supports yes/no, multiple-choice and rating options, plus rubric-based grading of AI responses and agent actions.

    • Base model: a post-trained Qwen3.5-9B, with Microsoft saying it plans to rebase the approach on other models, including Microsoft AI (MAI) and OpenAI models.
    • Accuracy: 83.5% average across 36 benchmarks and 147,137 held-out questions, ahead of the next best model in Microsoft's chart (Quyet-1.0-Large at 81.9%).
    • Speed: 85 ms median (P50) and 125 ms P95 through Foundry, against 3.01 s for GPT-6 Sol. Microsoft rounds the ratio to 35x.
    • Calibration: 92.2 on Microsoft's calibration scale, second to Quyet-1.0-Large at 93.1.
    • Price: $0.042 per million input tokens, free output tokens.
    • Availability: Microsoft Foundry now; OpenRouter listed by Microsoft, with some trade reports describing it as coming soon.
    • Caveat: this is a vendor launch. Per-benchmark scores, context length, rate limits and the calibration method are not fully published, and no independent replication has appeared yet.

    Our view: the idea is sound and overdue. Treat the model as a fast gate, not a brain, and run it against your own labelled examples before you trust a number from a launch post.

    What Is a Decision Model?

    A normal large language model answers by generating tokens one at a time. Ask it "Is this support ticket urgent?" and it may produce a paragraph of reasoning before it commits to "yes". That is flexible, but it is slow, it costs money for every token, and the answer can wobble if you rephrase the question or reorder the options.

    A decision model removes the generation step. It reads the request once, scores each allowed option, and returns the scores as probabilities. There is no free text to parse and no chance of the model inventing a fourth option. For software this is a much better interface: a number between 0 and 1 can be compared with a threshold, logged, audited and fed to a rule.

    This is not a brand-new idea. Classifiers and reward models have worked this way for years, and we covered a close cousin in our review of Liquid AI d1, the zero-token decision model. What is new is a hyperscaler putting a general-purpose version in its main model catalogue, with a published evaluation, an enterprise API and a stated roadmap to retrain it on its own and OpenAI models.

    Microsoft lists a wide set of intended jobs: agent controls, model routing, skill-based decisions, data labelling, AI judging, intent analysis, incident routing, data validation, recommendations, search relevance, content filtering, code scanning, safety and security screening, computer and UI use, robotics and scientific discovery. Most of those are really one task wearing different clothes: pick one of N, and say how sure you are.

    How It Works

    Three design choices matter. First, the closed option set. Because the model only scores options you provide, the output is always valid for your application. Second, single-pass scoring. The probabilities come from one forward pass rather than a sampled chain of thought, which is where most of the latency saving comes from. Third, calibration as a product feature. The probability is exposed in the API so an application can decide when to act, defer or ask for review.

    Microsoft states the calibration goal plainly: a 90% prediction should be correct about nine times in ten on representative cases. That is the right target, because an uncalibrated score is just a vibe with decimals. Note the wording, though. On Microsoft's launch page calibration is described as a goal; the measured calibration figures only appear in the comparison chart that was added afterwards, which we cover below.

    The model is built on Qwen3.5-9B, an open-weight model from Alibaba. Starting from a small open base and post-training it for one narrow interface is a pattern we have seen across the industry, from agent harness models to the Qwen 4 family now in training. The difference here is the interface: instead of making the small model a worse chatbot, Microsoft has made it a better scorer.

    In practice a call looks like a classification request with a rubric. You describe the decision ("Should this agent action be allowed given the policy below?"), list the options, and read back probabilities. Microsoft says it also supports grading, where the options are score levels on a rubric and the "response" under review is an AI answer or an agent action. That makes it a drop-in replacement for the "LLM as judge" pattern in evaluation pipelines.

    Benchmarks: What Microsoft Reports

    Microsoft says Decision-1 achieved the highest accuracy in a 36-benchmark evaluation covering nearly 150,000 questions that were withheld from training. The comparison chart puts the exact total at 147,137 questions, with the benchmark set spanning routing, ranking, long context, multilingual, out-of-distribution, reasoning and safety tasks. Microsoft also says it tested against top models on JevBench plus 36 additional public and private benchmarks.

    Microsoft chart comparing average accuracy, median latency and calibration of Microsoft-Decision-1 against Quyet-1.0-Large, Surogate Rune 26B-A4B, GPT-6 Luna Decisions, deck-31B, H2O-Lightning-4B, Strands-Decider 2B, Jev 1.13.0 and GPT-6 Sol
    Microsoft's "Accuracy, latency, and calibration, side by side" chart for Microsoft-Decision-1. Accuracy is the mean over 36 benchmarks (147,137 questions); latency is the JevBench v1.6.1 adjusted median checked on 07/10/2026. Source: Microsoft (Command Line), as reproduced by TestingCatalog.

    Here are the figures read directly from that chart:

    ModelAvg accuracyMedian latencyCalibration (100 = perfect)
    Microsoft-Decision-183.5%85 ms (P95 125 ms)92.2
    Quyet-1.0-Large81.9%380 ms93.1
    Surogate Rune 26B-A4B79.7%380 ms91.8
    GPT-6 Luna Decisions79.4%300 ms89.9
    deck-31B77.8%400 ms83.5
    H2O-Lightning-4B v1.177.2%210 ms91.8
    Strands-Decider 2B (AWS)54.8% on 23 of 36 benchmarksnot measurednot scored
    Jev 1.13.0not ranked240 msnot scored
    GPT-6 Sol (reference)not ranked3.01 snot scored

    Three things stand out. The accuracy lead is real but not enormous: 1.6 points over the runner-up. The latency lead is the headline: 85 ms is 4.5 times quicker than Quyet-1.0-Large (380 ms), roughly 2.5 times quicker than the fastest rival in the chart (H2O-Lightning-4B at 210 ms), and 35 times quicker than GPT-6 Sol. And on calibration Decision-1 is second, 0.9 behind Quyet, which the chart itself admits.

    The footnotes matter. Latency for Decision-1 was measured through Foundry in the same region, while the others come from the JevBench leaderboard's adjusted medians. The chart also marks which models rank on the JevBench leaderboard (Quyet-1.0-Large is first on it, as of 08/10/2026). Strands-Decider 2B only answered 23 of the 36 benchmarks, so its 54.8% is not like-for-like.

    Be careful with the 35x number. It compares a model that scores fixed options with one that generates text. GPT-6 Sol is a general-purpose frontier model, and 3.01 s reflects it doing a much bigger job. A fair reading is that for one narrow task, a purpose-built 9B scorer is more than an order of magnitude faster than asking a frontier model, which is expected. The interesting part is that it is also at least as accurate on Microsoft's chosen benchmarks.

    Robustness and Safety Testing

    A decision model is only useful if it gives the same answer when the question is asked slightly differently. Microsoft says it perturbed requests in eight ways and the decision flipped in 1.3% of cases on average. When option descriptions were paraphrased, or options were reversed or shuffled, there were zero flips. That addresses a well-known weakness of LLM judges, which can favour the first option or change their mind when the order changes.

    On safety, Microsoft reports testing on 5,250 requests across 11 benchmarks covering harmful content, jailbreaks and prompt injection, and says the model refused harmful behaviour while keeping high utility. Treat that as a headline only: the launch page does not list per-benchmark results, and "refused" is an odd verb for a model that outputs probabilities, so we would want to see how refusals are represented in the API.

    Prompt injection deserves its own warning. Because a decision model sits in the control path of an agent, an attacker who can influence its input can try to flip a gate from "block" to "allow". Low flip rates under paraphrase say little about adversarial text, so keep deterministic checks behind any high-stakes gate.

    Internal Trials: Xbox, Copilot and Incident Response

    Microsoft shares four internal results, all of them self-reported:

    • Xbox Research sorted more than 10,000 open-ended feedback items into fixed themes. Quality was competitive with GPT-6 Sol, while running over 14 times faster and 200 times cheaper.
    • The Copilot team used it for quality control of chat and agentic responses, finding quality competitive with GPT-5.6 Luna and 100 times faster.
    • Incident response saw it perform better and faster than an LLM at retrieving knowledge during live incidents.
    • Microsoft Discovery used it for adaptive replanning, where its scoring was 46 times more consistent than an LLM-based score at about three times the speed. Trade coverage rounds this to "nearly 4x faster", so the exact multiple is unclear.

    These are the right kinds of workload: high volume, repetitive, bounded. They are also the kinds of workload where the cheapest option wins on cost alone, so the "competitive quality" claim carries the weight. Note that the Xbox result is a labelling task against a frontier model's labels, which tells you about agreement with that model, not necessarily about ground truth.

    Why Agent Builders Should Care

    Agentic systems make dozens of small decisions per task. Which tool should run next? Is the output good enough, or should the step retry? Does this action need human approval? Which model should handle this request? Today those decisions are usually handled by prompting the same large model that does the work, which adds latency and cost to every step and makes the control logic as unpredictable as the model itself.

    A decision model offers a different architecture: a small, cheap, calibrated referee that sits beside the worker. Microsoft pitches it as exactly that, a control layer where the confidence score determines whether software acts, defers, retries, escalates or hands the work to a model, tool or person. We have seen the demand for this pattern in always-on assistants such as Microsoft Copilot Autopilot, where every autonomous action raises the question of who checks it, and in browser agents like Hark Handoff, where the hard part is deciding when to stop and ask a human.

    Model routing is the clearest cost lever. If a cheap scorer can decide that a request needs a frontier model, you pay frontier prices only for those requests. Pricing context: GPT-6 Sol and Luna were repriced earlier this season (see our GPT-6 Sol and Luna review), but even a discounted frontier call is far more expensive than a $0.042-per-million-token scorer with free output.

    The same pattern applies to evaluation. Teams that grade agent transcripts with a frontier model pay for every judgement. A fast rubric scorer lets you grade every step of every run, not a sample, which matters for the safety work we discuss in Anthropic's multiagent safety research, where the failures that count are rare and easy to miss in a sample.

    Pricing and Availability

    Microsoft lists $0.042 per million input tokens and no charge for output tokens. In pounds that is roughly 3p per million input tokens at current rates, and because the output is a handful of probability values, the effective cost per decision is tiny. The model is available in Microsoft Foundry now, and Microsoft's launch post says it is also available through OpenRouter; some trade reports said OpenRouter support was still "coming soon", so check the catalogue before you plan around it. Documentation is on Microsoft Learn.

    What Microsoft has not published is also useful to know: the context window, rate limits, the full benchmark list with per-benchmark scores, and the method behind the calibration score. If your inputs are long documents rather than short requests, ask for the context limit before designing around it.

    How It Compares

    Against LLM judges. Decision-1 trades explanation for speed. You get a probability and no rationale. For debugging you may still want a larger model to explain disagreements, but for the bulk of grading the scorer is enough.

    Against other decision models. On Microsoft's chart it leads on accuracy and latency, trails Quyet-1.0-Large on calibration, and shows that Liquid AI, H2O, Surogate and others are already competing here. Our Liquid AI d1 write-up covers a rival with a similar zero-token pitch; the two have not been compared head to head in any source we could find.

    Against the frontier models. Frontier models still win wherever reasoning in the open matters. Decision-1 does not browse, code or plan. Think of it as the traffic light, not the driver.

    Against doing nothing. Many teams currently use regular expressions, keyword rules or a frontier model prompt for these decisions. A calibrated scorer is a clear upgrade over rules for fuzzy cases and a clear saving over frontier prompts for high-volume ones.

    Limitations and Open Questions

    • Vendor-only evidence. The accuracy, latency, calibration and internal-trial figures all come from Microsoft. No independent leaderboard run has been published.
    • Benchmark mix unknown. The launch page does not itemise the 36 benchmarks, so we cannot tell how much of the suite resembles your workload.
    • Two different speed ratios. Reports quote 4.5x faster than the runner-up (from the chart) and 2.5x faster than the fastest rival (from the post text). They are consistent only once you know which rival each refers to.
    • Latency comparisons are not like-for-like. Decision-1 was timed through Foundry; others use JevBench adjusted medians.
    • Calibration is second-best. It is a headline feature, yet Quyet-1.0-Large scores higher on Microsoft's own chart.
    • Closed options only. It cannot answer questions you did not anticipate.
    • Control-path risk. A gate model is a target for prompt injection.
    • Base-model licence. It inherits Qwen3.5-9B, so check data-handling and licence terms for regulated workloads.

    Who Should Use It

    Good fit: teams running high-volume classification, routing or moderation; agent builders who want a cheap approval gate; evaluation teams grading transcripts at scale; anyone paying a frontier model to emit one-word answers.

    Poor fit: open-ended generation, anything needing an explanation, workloads with very long inputs (until the context limit is published), and safety-critical gates with no deterministic backstop.

    How to trial it: collect 200 to 500 examples you have already labelled, run Decision-1 and your current approach side by side, compare accuracy and the calibration curve (do 90% predictions come true about 90% of the time?), then set thresholds that route uncertain cases to a larger model or a human.

    The Bottom Line

    Microsoft-Decision-1 is a small model with a clear job and a sensible interface. The numbers on Microsoft's own chart, 83.5% accuracy at 85 ms, would make it a very attractive control layer if they hold up on your data. The big claim, 35 times faster than GPT-6 Sol, is true in the narrow sense that scoring is cheaper than generating, and it should not be read as "a 9B model beats GPT-6 Sol". Test it, keep a deterministic backstop for anything risky, and watch for independent evaluations. If you want the wider context for where this fits in the week's news, see our write-ups of the Gemini 4 Carbon leak and Gemini 4 Argon, which dominated the same news cycle.

    Sources

    Last updated: 11/10/2026. Sourced from Microsoft's launch post and trade coverage. We have not tested Decision-1 ourselves; every performance figure is vendor-reported.

    Free Guide

    Get the free guide: Claude vs ChatGPT, Gemini & Grok

    A 20-page playbook covering everything you need to choose and use the big four AI models in 2026, full cost and feature comparisons, what each is best (and worst) at, and how-tos for images, vectors, building a website, Claude Code and more.

    Pop your email in to get it free
    Preview of the free guide: Claude vs ChatGPT, Gemini and Grok, 2026 features, pricing and what-you-can-do comparison.

    Frequently Asked Questions

    What is Microsoft-Decision-1?
    Microsoft-Decision-1 is a decision-scoring model launched on 09/10/2026 in Microsoft Foundry. You give it a fixed set of options (yes or no, multiple choice, or a rating scale) and it returns a calibrated probability for each in a single pass, rather than generating free text. Microsoft says it is a post-trained version of Alibaba's open-weight Qwen3.5-9B and is designed for routing, classification, verification and agent control.
    How fast is Microsoft-Decision-1 compared with GPT-6 Sol?
    Microsoft says it is about 35 times faster than GPT-6 Sol at median (P50) latency. Microsoft's own comparison chart shows 85 milliseconds for Decision-1 against 3.01 seconds for GPT-6 Sol, with a P95 of 125 milliseconds for Decision-1 measured through Foundry. These are vendor-reported figures and the two models are doing different jobs: one scores fixed options, the other generates text.
    How much does Microsoft-Decision-1 cost?
    Microsoft lists $0.042 per million input tokens (roughly 3p) and no charge for output tokens. It is available in Microsoft Foundry and, according to Microsoft, through OpenRouter, with documentation on Microsoft Learn. Microsoft has not published a context window or rate limits on its launch page.
    Is Decision-1 the same as an LLM judge?
    It can do the job of an LLM judge, since Microsoft says it supports rubric-based grading of AI responses and agent actions, but it is built differently. A normal LLM judge generates reasoning and a verdict token by token. Decision-1 reads the request and outputs probabilities for the allowed options in one pass, which is why it is much faster and cheaper but also cannot explain its reasoning in prose.
    Should I replace GPT-class models with Decision-1?
    No. It is a control-layer model for bounded choices such as routing, labelling, filtering and gating agent actions. It does not write, reason in the open or use tools. The sensible pattern is to place it in front of or beside a larger model, use its probability to decide when to act, defer or escalate, and validate it on your own labelled data first, because most published numbers are Microsoft's own.

    Explore more AI tool comparisons

    In-depth reviews, benchmarks and guides to help you choose the right AI tools.

    Browse all reviews
    AI Tools Review Editorial Team

    AI Tools Review Editorial Team Expert verified

    Our editorial team consists of veteran AI researchers, software engineers, and industry analysts. We spend hundreds of hours benchmarking frontier models natively to provide you with objective, actionable intelligence on agentic AI capabilities and cybersecurity landscapes.