Key Points
- Microsoft released Microsoft-Decision-1, a fast scoring model for routing, classification and workflow control.
- Input costs 0.042 dollars per million tokens, with output tokens free, Microsoft said.
- It targets the cost of calling large models for every routine agent decision.
The latest:
A new Microsoft model that does nothing but score decisions went live on October 9, 2026, available in Microsoft Foundry and through OpenRouter. Microsoft-Decision-1 returns a calibrated probability for each option in a fixed set through a structured API call, covering yes/no, multiple choice and rating-scale questions. Microsoft priced input at 0.042 dollars per million tokens and made output tokens free.
Details:
- The architecture: Microsoft said Decision-1 is built on Qwen3.5-9B with post-training aimed at single-pass decision scoring rather than open-ended generation. The company said it plans to rebuild the model soon on other bases, including its own Microsoft AI (MAI) models and OpenAI models, without announcing a timeline.
- The accuracy claim: According to Microsoft, Decision-1 posted the highest accuracy in a comparison spanning 36 benchmarks and roughly 150,000 questions, and the company said those benchmarks were held out of training. It also said it tested several models from the top of the JevBench leaderboard across 36 additional public and private benchmarks, with Decision-1 ranking best.
- The speed numbers: Microsoft said the model was the fastest it measured: 2.5 times faster than runner-up H2O-Lightning-4B v1.1 and 35 times faster than GPT-6 Sol. P50 latency was also reported at roughly 35 times faster than GPT-6 Sol.
- Robustness testing: Across 8 perturbation types, decisions flipped in 1.3% of cases on average, Microsoft said, and zero flips occurred when option descriptions were rephrased, reversed or shuffled. The company described calibration in principle, saying a 90% prediction should be right about 9 times out of 10, without publishing a measured figure.
- Safety evaluation: Microsoft said it ran 5,250 prompts across 11 benchmarks covering harmful content, jailbreaking and prompt injection. The company said the model refused harmful behavior while keeping high usefulness, and did not break out refusal rates by category.
- Internal deployments: The Xbox research team classified more than 10,000 player feedback items at quality Microsoft called competitive with GPT-6 Sol, more than 14 times faster and 200 times cheaper. The Copilot team used it to monitor chat and agent response quality at quality competitive with GPT5.6 Luna and 100 times the speed.
- Other results: In Microsoft Discovery, the company said the model scored 46 times higher consistency on adaptive replanning than an evaluation built on a large language model, at three times the speed. For incident response, Microsoft reported better and faster knowledge retrieval than a large language model without releasing figures.
- The comparison set: The appendix lists the models benchmarked against Decision-1: Quyet-1.0-Large, Surogate Rune 26B-A4B, GPT-6 Luna Decisions, deck-31B, H2O-Lightning-4B and Strands-Decider 2B. Documentation sits on Microsoft Learn, alongside two demos comparing the model with GPT-6 Sol on query classification and a backpack purchase.
- Intended uses: Microsoft listed agent guardrails, model routing, skill-based decisions, data labeling, automated judging, intent analysis, incident routing, data validation, recommendations, search relevance, content filtering, code review, security triage, computer and interface use, robotics and scientific discovery.
Background:
The blog post, written by Achint Srivastava on Microsoft’s Command Line blog, carried an editor’s note saying it was updated to add JevBench accuracy and calibration results.
Between the lines:
Every headline figure here is Microsoft’s own, measured by Microsoft, including the benchmark holdout claim. The pricing structure points at the real target: by charging 0.042 dollars per million input tokens and nothing for output, the company is positioning Decision-1 as the layer enterprises call thousands of times per workflow, reserving expensive frontier models for the steps that need generation.
What’s next
Watch for independent benchmark results against the listed comparison models, a published measured calibration figure, and whether Microsoft ships the promised rebuilds on MAI and OpenAI bases.