Microsoft-Decision-1: a 9B scorer that wins Microsoft's own scorecard
Microsoft post-trained Qwen3.5-9B into a 9B decision-scoring model, then beat its own benchmark field. We separate the shipped artifact from the category play.

Microsoft introduced Microsoft-Decision-1 on 9 October, a 9B model that does not write prose. Hand it a fixed set of options and it returns a calibrated probability for each one, in a single pass, through a structured API call. The company's announcement presents it as the arrival of decision models as a category beside large language models and agents. The headline numbers are strong: 83.5% mean accuracy across 36 benchmarks, a median latency of 85 milliseconds, calibration Microsoft ranks third in the field. Every one of those figures comes from Microsoft, about a model Microsoft built, measured against a field Microsoft chose.
A scorecard built and won in the same building
Microsoft's comparison places its model first on accuracy and first on speed, and the piece is worth reading as an assembly: 36 benchmarks spanning 147,137 questions the company describes as kept blind from training, plus a rival set Microsoft picked. The rows are third-party models; the scoring is not.
| Model | Mean accuracy | Median latency | Calibration |
|---|---|---|---|
| Microsoft-Decision-1 | 83.5% | 85 ms | 92.2 |
| Jev 1.13.0 | 82.3% | 240 ms | 93.7 |
| Quyet-1.0-Large | 81.9% | 380 ms | 93.1 |
| Surogate Rune 26B-A4B | 79.7% | 380 ms | 91.8 |
| GPT-6 Luna Decisions | 79.4% | 300 ms | 89.9 |
| H2O-Lightning-4B v1.1 | 77.2% | 210 ms | 91.8 |
The field itself is Microsoft's shortlist, drawn from the public JevBench leaderboard on 8 October, and it includes one entry the vendor footnotes away: Strands-Decider 2B answered only 23 of the 36 benchmarks and was left unranked.
Two things stand out. On calibration, where confidence should match how often the model is right, Microsoft places its own entry third, behind Jev 1.13.0 and Quyet-1.0-Large, and says so in the chart. And the accuracy column is tight: less than three points separate the runner-up from GPT-6 Luna Decisions, five places down. The gap Microsoft is selling sits at the top of the latency column, not the accuracy one. The post also carries an editor's note that it was updated after publication to add benchmarks for Jev on accuracy and calibration, which is the kind of housekeeping a vendor does when the first table looked worse.
Somebody else's base model, restated as a new kind of model
Microsoft is candid about the mechanism. It post-trained Qwen3.5-9B for fast single-pass decision scoring and says it will "soon rebase it on other models, including Microsoft AI (MAI) and OpenAI." The capability is not a new architecture; it is a 9B open-weight base, steered toward a narrow task. That disclosure is refreshing next to vendors who call a fine-tune a new family, and it also bounds the claim. The engineering is what Microsoft did to the task, not to the network: it fixed the option types (yes/no, multiple choice, ratings, and rubric grading of AI responses and agent actions), trained for calibration rather than fluency, and tested whether equivalent inputs produce equivalent decisions. Microsoft reports the model flips on 1.3% of perturbations, with no flips when option descriptions are paraphrased or choices are reordered.
85 milliseconds is the whole argument
The latency case is stated cleanly enough to check. Adding 100 ms to each of 20 sequential decisions adds two seconds to a workflow, Microsoft writes, and it measures 85 ms at the median and 125 ms at p95. That is 2.5 times faster than the runner-up, H2O-Lightning-4B v1.1 at 210 ms, and 35 times faster than GPT-6 Sol at 3.01 seconds. The measurement comes from calls through Microsoft Foundry in the same region, which is a fair test and a narrow one, since it excludes the network hop most deployments will pay. On price, the same blog lists $0.042 per million input tokens with output free, and the free output is literal: there is no text to generate because the answer is read off the probability distribution rather than produced token by token.
The category is outrunning the model
Microsoft's real pitch is the label, and the label is not Microsoft's alone. It calls decision models a new category in AI, and it is one of several companies saying so at once. The Verge noted the launch landed days after OpenAI highlighted a Decisions API in public beta. Liquid AI shipped open-weight decision models the same day, and Cloudflare's Clef models arrived the week before. A category that three vendors enter inside two weeks is a product line forming, not a leaderboard win. Microsoft frames the buyer as an agent builder who needs to route, classify, prioritize and verify at machine speed, and OpenAI and the open-model vendors use the same language. The overlap is the point. For a buyer, that is the useful signal: the fastest decision scorer is worth less than a shortlist of them, because the thing being sold is a slot in an agent's loop, and slots get commoditized.
What would settle it
Three checks would move this from a vendor result to a fact. An independent evaluation on the routing, classification and scoring tasks Microsoft names, since the blinded 36-benchmark set is private and nobody outside can reproduce the 83.5%. A third-party calibration audit, because calibration is the claim that decides whether a system can act on the score without a human in the loop, and Microsoft already ranks itself third on its own measure. And latency measured end-to-end from a client outside Microsoft's region, because a same-region Foundry call is the best case rather than the typical one. Until one of those lands, the honest summary is that Microsoft shipped a fast, cheap, well-instrumented classifier, wrapped the launch in a category it is helping to define, and stacked the deck exactly where every vendor stacks it, on a benchmark set of its own making.


