Quick Answer:
Microsoft-Decision-1 is a small 9-billion-parameter decision model (a post-trained Qwen3.5-9B) that Microsoft launched in Foundry on 09/10/2026. Instead of writing text, it takes a fixed set of options and returns a calibrated probability for each in one pass. Microsoft reports the highest average accuracy across 36 benchmarks (83.5%), a median latency of 85 ms against 3.01 s for GPT-6 Sol (about 35x faster), and a price of $0.042 per million input tokens with free output. Almost every number is Microsoft's own, so treat it as a strong, unverified claim.
The most interesting AI model of the week is not another frontier chatbot. It is a small model that refuses to talk. Microsoft-Decision-1 does one thing: you hand it a question and a list of allowed answers, and it tells you how likely each answer is, in about a tenth of a second.
That sounds modest until you count how many decisions an AI agent makes per task. Every route, retry, approval, label and safety check is a decision, and today most of them are made by a large language model that generates hundreds of tokens to reach a one-word answer. Microsoft is betting that this layer deserves its own cheap, fast, calibrated model.
AI Revolution X covers Microsoft Decision-1 from the 7:39 chapter, after the Gemini 4 Carbon leak.
Executive Summary
What it is: Microsoft-Decision-1 is a "decision model". Microsoft describes decision models as purpose-built to deliver structured outputs that software can act on immediately. You supply a prompt and a closed set of options; the model returns a probability for each option through a structured API call. It supports yes/no, multiple-choice and rating options, plus rubric-based grading of AI responses and agent actions.
- Base model: a post-trained Qwen3.5-9B, with Microsoft saying it plans to rebase the approach on other models, including Microsoft AI (MAI) and OpenAI models.
- Accuracy: 83.5% average across 36 benchmarks and 147,137 held-out questions, ahead of the next best model in Microsoft's chart (Quyet-1.0-Large at 81.9%).
- Speed: 85 ms median (P50) and 125 ms P95 through Foundry, against 3.01 s for GPT-6 Sol. Microsoft rounds the ratio to 35x.
- Calibration: 92.2 on Microsoft's calibration scale, second to Quyet-1.0-Large at 93.1.
- Price: $0.042 per million input tokens, free output tokens.
- Availability: Microsoft Foundry now; OpenRouter listed by Microsoft, with some trade reports describing it as coming soon.
- Caveat: this is a vendor launch. Per-benchmark scores, context length, rate limits and the calibration method are not fully published, and no independent replication has appeared yet.
Our view: the idea is sound and overdue. Treat the model as a fast gate, not a brain, and run it against your own labelled examples before you trust a number from a launch post.
What Is a Decision Model?
A normal large language model answers by generating tokens one at a time. Ask it "Is this support ticket urgent?" and it may produce a paragraph of reasoning before it commits to "yes". That is flexible, but it is slow, it costs money for every token, and the answer can wobble if you rephrase the question or reorder the options.
A decision model removes the generation step. It reads the request once, scores each allowed option, and returns the scores as probabilities. There is no free text to parse and no chance of the model inventing a fourth option. For software this is a much better interface: a number between 0 and 1 can be compared with a threshold, logged, audited and fed to a rule.
This is not a brand-new idea. Classifiers and reward models have worked this way for years, and we covered a close cousin in our review of Liquid AI d1, the zero-token decision model. What is new is a hyperscaler putting a general-purpose version in its main model catalogue, with a published evaluation, an enterprise API and a stated roadmap to retrain it on its own and OpenAI models.
Microsoft lists a wide set of intended jobs: agent controls, model routing, skill-based decisions, data labelling, AI judging, intent analysis, incident routing, data validation, recommendations, search relevance, content filtering, code scanning, safety and security screening, computer and UI use, robotics and scientific discovery. Most of those are really one task wearing different clothes: pick one of N, and say how sure you are.
How It Works
Three design choices matter. First, the closed option set. Because the model only scores options you provide, the output is always valid for your application. Second, single-pass scoring. The probabilities come from one forward pass rather than a sampled chain of thought, which is where most of the latency saving comes from. Third, calibration as a product feature. The probability is exposed in the API so an application can decide when to act, defer or ask for review.
Microsoft states the calibration goal plainly: a 90% prediction should be correct about nine times in ten on representative cases. That is the right target, because an uncalibrated score is just a vibe with decimals. Note the wording, though. On Microsoft's launch page calibration is described as a goal; the measured calibration figures only appear in the comparison chart that was added afterwards, which we cover below.
The model is built on Qwen3.5-9B, an open-weight model from Alibaba. Starting from a small open base and post-training it for one narrow interface is a pattern we have seen across the industry, from agent harness models to the Qwen 4 family now in training. The difference here is the interface: instead of making the small model a worse chatbot, Microsoft has made it a better scorer.
In practice a call looks like a classification request with a rubric. You describe the decision ("Should this agent action be allowed given the policy below?"), list the options, and read back probabilities. Microsoft says it also supports grading, where the options are score levels on a rubric and the "response" under review is an AI answer or an agent action. That makes it a drop-in replacement for the "LLM as judge" pattern in evaluation pipelines.
Benchmarks: What Microsoft Reports
Microsoft says Decision-1 achieved the highest accuracy in a 36-benchmark evaluation covering nearly 150,000 questions that were withheld from training. The comparison chart puts the exact total at 147,137 questions, with the benchmark set spanning routing, ranking, long context, multilingual, out-of-distribution, reasoning and safety tasks. Microsoft also says it tested against top models on JevBench plus 36 additional public and private benchmarks.

Here are the figures read directly from that chart:
| Model | Avg accuracy | Median latency | Calibration (100 = perfect) |
|---|---|---|---|
| Microsoft-Decision-1 | 83.5% | 85 ms (P95 125 ms) | 92.2 |
| Quyet-1.0-Large | 81.9% | 380 ms | 93.1 |
| Surogate Rune 26B-A4B | 79.7% | 380 ms | 91.8 |
| GPT-6 Luna Decisions | 79.4% | 300 ms | 89.9 |
| deck-31B | 77.8% | 400 ms | 83.5 |
| H2O-Lightning-4B v1.1 | 77.2% | 210 ms | 91.8 |
| Strands-Decider 2B (AWS) | 54.8% on 23 of 36 benchmarks | not measured | not scored |
| Jev 1.13.0 | not ranked | 240 ms | not scored |
| GPT-6 Sol (reference) | not ranked | 3.01 s | not scored |
Three things stand out. The accuracy lead is real but not enormous: 1.6 points over the runner-up. The latency lead is the headline: 85 ms is 4.5 times quicker than Quyet-1.0-Large (380 ms), roughly 2.5 times quicker than the fastest rival in the chart (H2O-Lightning-4B at 210 ms), and 35 times quicker than GPT-6 Sol. And on calibration Decision-1 is second, 0.9 behind Quyet, which the chart itself admits.
The footnotes matter. Latency for Decision-1 was measured through Foundry in the same region, while the others come from the JevBench leaderboard's adjusted medians. The chart also marks which models rank on the JevBench leaderboard (Quyet-1.0-Large is first on it, as of 08/10/2026). Strands-Decider 2B only answered 23 of the 36 benchmarks, so its 54.8% is not like-for-like.
Be careful with the 35x number. It compares a model that scores fixed options with one that generates text. GPT-6 Sol is a general-purpose frontier model, and 3.01 s reflects it doing a much bigger job. A fair reading is that for one narrow task, a purpose-built 9B scorer is more than an order of magnitude faster than asking a frontier model, which is expected. The interesting part is that it is also at least as accurate on Microsoft's chosen benchmarks.
Robustness and Safety Testing
A decision model is only useful if it gives the same answer when the question is asked slightly differently. Microsoft says it perturbed requests in eight ways and the decision flipped in 1.3% of cases on average. When option descriptions were paraphrased, or options were reversed or shuffled, there were zero flips. That addresses a well-known weakness of LLM judges, which can favour the first option or change their mind when the order changes.
On safety, Microsoft reports testing on 5,250 requests across 11 benchmarks covering harmful content, jailbreaks and prompt injection, and says the model refused harmful behaviour while keeping high utility. Treat that as a headline only: the launch page does not list per-benchmark results, and "refused" is an odd verb for a model that outputs probabilities, so we would want to see how refusals are represented in the API.
Prompt injection deserves its own warning. Because a decision model sits in the control path of an agent, an attacker who can influence its input can try to flip a gate from "block" to "allow". Low flip rates under paraphrase say little about adversarial text, so keep deterministic checks behind any high-stakes gate.
Internal Trials: Xbox, Copilot and Incident Response
Microsoft shares four internal results, all of them self-reported:
- Xbox Research sorted more than 10,000 open-ended feedback items into fixed themes. Quality was competitive with GPT-6 Sol, while running over 14 times faster and 200 times cheaper.
- The Copilot team used it for quality control of chat and agentic responses, finding quality competitive with GPT-5.6 Luna and 100 times faster.
- Incident response saw it perform better and faster than an LLM at retrieving knowledge during live incidents.
- Microsoft Discovery used it for adaptive replanning, where its scoring was 46 times more consistent than an LLM-based score at about three times the speed. Trade coverage rounds this to "nearly 4x faster", so the exact multiple is unclear.
These are the right kinds of workload: high volume, repetitive, bounded. They are also the kinds of workload where the cheapest option wins on cost alone, so the "competitive quality" claim carries the weight. Note that the Xbox result is a labelling task against a frontier model's labels, which tells you about agreement with that model, not necessarily about ground truth.
Why Agent Builders Should Care
Agentic systems make dozens of small decisions per task. Which tool should run next? Is the output good enough, or should the step retry? Does this action need human approval? Which model should handle this request? Today those decisions are usually handled by prompting the same large model that does the work, which adds latency and cost to every step and makes the control logic as unpredictable as the model itself.
A decision model offers a different architecture: a small, cheap, calibrated referee that sits beside the worker. Microsoft pitches it as exactly that, a control layer where the confidence score determines whether software acts, defers, retries, escalates or hands the work to a model, tool or person. We have seen the demand for this pattern in always-on assistants such as Microsoft Copilot Autopilot, where every autonomous action raises the question of who checks it, and in browser agents like Hark Handoff, where the hard part is deciding when to stop and ask a human.
Model routing is the clearest cost lever. If a cheap scorer can decide that a request needs a frontier model, you pay frontier prices only for those requests. Pricing context: GPT-6 Sol and Luna were repriced earlier this season (see our GPT-6 Sol and Luna review), but even a discounted frontier call is far more expensive than a $0.042-per-million-token scorer with free output.
The same pattern applies to evaluation. Teams that grade agent transcripts with a frontier model pay for every judgement. A fast rubric scorer lets you grade every step of every run, not a sample, which matters for the safety work we discuss in Anthropic's multiagent safety research, where the failures that count are rare and easy to miss in a sample.
Pricing and Availability
Microsoft lists $0.042 per million input tokens and no charge for output tokens. In pounds that is roughly 3p per million input tokens at current rates, and because the output is a handful of probability values, the effective cost per decision is tiny. The model is available in Microsoft Foundry now, and Microsoft's launch post says it is also available through OpenRouter; some trade reports said OpenRouter support was still "coming soon", so check the catalogue before you plan around it. Documentation is on Microsoft Learn.
What Microsoft has not published is also useful to know: the context window, rate limits, the full benchmark list with per-benchmark scores, and the method behind the calibration score. If your inputs are long documents rather than short requests, ask for the context limit before designing around it.
How It Compares
Against LLM judges. Decision-1 trades explanation for speed. You get a probability and no rationale. For debugging you may still want a larger model to explain disagreements, but for the bulk of grading the scorer is enough.
Against other decision models. On Microsoft's chart it leads on accuracy and latency, trails Quyet-1.0-Large on calibration, and shows that Liquid AI, H2O, Surogate and others are already competing here. Our Liquid AI d1 write-up covers a rival with a similar zero-token pitch; the two have not been compared head to head in any source we could find.
Against the frontier models. Frontier models still win wherever reasoning in the open matters. Decision-1 does not browse, code or plan. Think of it as the traffic light, not the driver.
Against doing nothing. Many teams currently use regular expressions, keyword rules or a frontier model prompt for these decisions. A calibrated scorer is a clear upgrade over rules for fuzzy cases and a clear saving over frontier prompts for high-volume ones.
Limitations and Open Questions
- Vendor-only evidence. The accuracy, latency, calibration and internal-trial figures all come from Microsoft. No independent leaderboard run has been published.
- Benchmark mix unknown. The launch page does not itemise the 36 benchmarks, so we cannot tell how much of the suite resembles your workload.
- Two different speed ratios. Reports quote 4.5x faster than the runner-up (from the chart) and 2.5x faster than the fastest rival (from the post text). They are consistent only once you know which rival each refers to.
- Latency comparisons are not like-for-like. Decision-1 was timed through Foundry; others use JevBench adjusted medians.
- Calibration is second-best. It is a headline feature, yet Quyet-1.0-Large scores higher on Microsoft's own chart.
- Closed options only. It cannot answer questions you did not anticipate.
- Control-path risk. A gate model is a target for prompt injection.
- Base-model licence. It inherits Qwen3.5-9B, so check data-handling and licence terms for regulated workloads.
Who Should Use It
Good fit: teams running high-volume classification, routing or moderation; agent builders who want a cheap approval gate; evaluation teams grading transcripts at scale; anyone paying a frontier model to emit one-word answers.
Poor fit: open-ended generation, anything needing an explanation, workloads with very long inputs (until the context limit is published), and safety-critical gates with no deterministic backstop.
How to trial it: collect 200 to 500 examples you have already labelled, run Decision-1 and your current approach side by side, compare accuracy and the calibration curve (do 90% predictions come true about 90% of the time?), then set thresholds that route uncertain cases to a larger model or a human.
The Bottom Line
Microsoft-Decision-1 is a small model with a clear job and a sensible interface. The numbers on Microsoft's own chart, 83.5% accuracy at 85 ms, would make it a very attractive control layer if they hold up on your data. The big claim, 35 times faster than GPT-6 Sol, is true in the narrow sense that scoring is cheaper than generating, and it should not be read as "a 9B model beats GPT-6 Sol". Test it, keep a deterministic backstop for anything risky, and watch for independent evaluations. If you want the wider context for where this fits in the week's news, see our write-ups of the Gemini 4 Carbon leak and Gemini 4 Argon, which dominated the same news cycle.
Sources
- Microsoft Command Line: Microsoft-Decision-1 in Foundry: Microsoft's launch post (accuracy, robustness, safety, internal trials, pricing, limits).
- TestingCatalog: Microsoft launches Decision-1 in Foundry (09/10/2026): independent summary, Satya Nadella's announcement post and the accuracy/latency/calibration chart.
- Runtime Wire: Microsoft puts a 9B decision model in Foundry for agent workflows (09/10/2026): context on Achint Srivastava and the production-trust caveat.
- AI Revolution X on YouTube: Google's New CARBON Is Apparently a MONSTER and WorldofAI: HUGE Gemini 4 Carbon LEAKS: creator coverage of the launch.
- Images: hero card and comparison chart are Microsoft launch graphics as reproduced by TestingCatalog; both are vendor-produced and not independent measurements.
Last updated: 11/10/2026. Sourced from Microsoft's launch post and trade coverage. We have not tested Decision-1 ourselves; every performance figure is vendor-reported.
Get the free guide: Claude vs ChatGPT, Gemini & Grok
A 20-page playbook covering everything you need to choose and use the big four AI models in 2026, full cost and feature comparisons, what each is best (and worst) at, and how-tos for images, vectors, building a website, Claude Code and more.





