
Microsoft’s launch artwork for Microsoft-Decision-1. Image: Image: Microsoft
Microsoft has launched Microsoft-Decision-1, an AI model that never writes a sentence. Give it a question and a fixed list of answers, and it hands back a probability for each one, in a single pass. Under the Microsoft name sits a model from China: Microsoft built it by post-training Alibaba’s open-weight Qwen3.5-9B.
The 9-billion-parameter model went live in Microsoft Foundry on October 9 and is also listed on OpenRouter. Every performance figure below comes from Microsoft’s own testing and hasn’t been independently checked.
A model that answers with a number
Most chatbots generate text and reason their way to an answer. Microsoft calls Decision-1 a “decision model”: it scores yes/no questions, multiple-choice options and ratings, and can grade an AI’s response or an agent’s proposed action against a rubric. The output is a set of probabilities that software can act on straight away, with no explanation attached.
Microsoft says the probabilities are calibrated, meaning a 90% answer should be right about nine times in ten. Apps can use that confidence to decide when to act, when to wait and when to ask a person.
The suggested uses are the plumbing of AI agents: deciding whether an agent should continue, stop, retry or hand off to a human; routing a request to the cheapest model that can handle it; labelling data; screening requests for safety risks; and picking the next click in computer-use tasks or the next move for a robot.
35 times faster than GPT-6 Sol, Microsoft says
Microsoft says Decision-1 had the highest accuracy in its own comparison across 36 benchmarks, nearly 150,000 questions kept out of training. Its median (P50) latency was about 35 times faster than GPT-6 Sol, and 2.5 times faster than the runner-up, H2O-Lightning-4B v1.1. The rivals it lists include OpenAI’s GPT-6 Luna Decisions and smaller open models such as Quyet-1.0-Large and Strands-Decider 2B. An editor’s note on the post says it was updated after launch to add accuracy and calibration results for Jev, which tops the JevBench leaderboard Microsoft drew rivals from.
On consistency, Microsoft rewords, reorders and reformats each request eight ways and checks whether the answer flips. It says Decision-1 changed its decision 1.3% of the time on average, and never when options were paraphrased, reversed or shuffled. It also tested 5,250 harmful, jailbreak and prompt-injection requests across 11 safety benchmarks.
Pricing is $0.042 per million input tokens, and output tokens are free.
“We’re already testing it across Microsoft”
Microsoft CEO Satya Nadella announced the model on X:
Introducing Microsoft-Decision-1, our new model for fast decision-making.
— Satya Nadella (@satyanadella) October 9, 2026
It delivers top performance on structured decision tasks, outperforming both LLMs and other decision models in latency and quality.
We’re already testing it across Microsoft for everything from incident response and quality control to scientific discovery.
Xbox Research used it to sort more than 10,000 pieces of player feedback from surveys, Steam and X into themes, and found it competitive with GPT-6 Sol on quality while running over 14 times faster at 200 times lower cost, Microsoft says. The Copilot team found it competitive with GPT-5.6 Luna at grading chat and agent responses, and 100 times faster. In Microsoft Discovery, its science platform, Microsoft says it scored 46 times more consistently than an LLM grader.
Microsoft says it will soon rebase the model on others, including its own MAI models and OpenAI’s. That follows its push this week to run smaller local models and sandboxed agents in Windows, and Nadella’s call a day later to treat every AI model as if it’s already been compromised.
Why it matters
As agents take more actions on their own, the cheap, fast calls that decide what they do next matter as much as the big model doing the work. That Microsoft, OpenAI’s biggest backer, built this one on an Alibaba model shows how far Chinese open-weight models have spread, even as few Chinese labs publish safety results for them. Until outside tests arrive, the speed and accuracy claims are Microsoft’s alone.
Sources: Microsoft Command Line; Microsoft Foundry model catalogue; Satya Nadella on X; Qwen3.5-9B on Hugging Face.


