Decision models are quickly emerging as an important new category in AI. Unlike LLMs, which are designed to generate text or reason through complex problems, decision models are purpose-built to deliver structured outputs that software can immediately act on. And once you understand that capability—making decisions and classifying things at very low cost with high performance—all kinds of useful tasks get unlocked.
Today we’re introducing Microsoft-Decision-1, our new model for fast decision-scoring, available in Microsoft Foundry and coming soon through OpenRouter. This model is designed for routing, classification, prioritization, verification, and workflow control, making it easier to incorporate decision intelligence into existing applications, agents, and workflows in a secure, trusted environment. Microsoft-Decision-1 delivers top performance in latency and quality on structured decision tasks to outperform both LLMs and other decision models.
Microsoft-Decision-1 achieved the highest accuracy in our 36-benchmark comparison, spanning nearly 150,000 questions across benchmarks kept blind from training. And in our benchmarking, it was the fastest measured: 4.5 times quicker than Quyet-1.0-Large, the runner-up, and 35 times quicker than GPT-6 Sol.
To build Microsoft-Decision-1, we post trained Qwen3.5-9B for fast, single-pass decision scoring and will soon rebase it on other models, including Microsoft AI (MAI) and OpenAI. When given a fixed set of answer options, Microsoft-Decision-1 provides a calibrated probability score for each option. The model supports yes/no, multiple-choice, and rating options, as well as rubric-based grading of AI responses and agent actions, all through a simple structured API call.
To build a reliable decision model, we had to address several challenges:
Each decision adds delay, especially when one step depends on another. For example, adding just 100 milliseconds to each of 20 sequential decisions adds two seconds to the overall workflow.
Microsoft-Decision-1 P50 latency is ~35x faster than GPT-6 Sol.
It’s easy to overfit a model for one benchmark or one type of decision task. We need to know whether that quality carries over to various tasks the model wasn’t trained on. That’s why we evaluated Microsoft-Decision-1 across dozens of benchmarks kept blinded from training, spanning routing, ranking, long context, multilingual and out-of-distribution tasks, reasoning, and safety. We also took several of the top public models on the popular open leaderboard JevBench and tested them across 36 additional public and private benchmarks. Microsoft-Decision-1 performed the best across these broader sets of benchmarks, demonstrating strong generalization.
Equivalent inputs should produce equivalent decisions. In production, states and instructions get paraphrased, option descriptions change, choices are reordered, keys change, and harmless formatting noise appears. None of those changes should materially alter the decision.
We perturb the same request in eight ways and measure how often the decision flips. Microsoft-Decision-1 changes its decision on 1.3% of perturbations on average with zero flips when option descriptions are paraphrased or when options are reversed or shuffled.
The probability itself is part of the API, not just a ranking score. Applications use confidence to decide when to act, defer, or ask for review, so a 90% prediction should be right about nine times out of 10 on representative cases.
A decision model should recognize harmful requests without needlessly blocking harmless ones. We tested Microsoft-Decision-1 on 5,250 requests across 11 benchmarks, covering harmful content, jailbreak attempts, and prompt injection, and found that the model successfully refused harmful behavior while retaining a high degree of utility.
Classification is a key use case for decision models. Check out how accurately and quickly Microsoft-Decision-1 can categorize a variety of queries compared to GPT-6 Sol:
Decision models can also be efficient for computer use scenarios. This demo shows how fast Microsoft-Decision-1 can complete the task of buying a backpack compared to GPT-6 Sol:
Here are some of the ways we’ve been testing Microsoft-Decision-1 internally, with a lot more to come.
XBOX Research used Microsoft-Decision-1 to process more than 10,000 open-ended pieces of feedback and reviews from surveys, STEAM, and Twitter/X and sort them into a fixed set of themes established by researchers to understand what people are saying about different games, launches, streams, and more. They found Microsoft-Decision-1 to be competitive on quality with GPT-6 Sol while running over 14 times faster and 200 times less expensive.
The Copilot team measures the quality of chat and agentic responses. Their testing found Microsoft-Decision-1 to be competitive with GPT5.6 Luna and 100 times faster.
Our on-call engineers use AI to retrieve relevant knowledge to respond to live incidents across logs, ticketing systems, calls, messages, and other data sources. Microsoft-Decision-1 performed better and faster than an LLM for knowledge retrieval.
Microsoft Discovery implements an adaptive replanning feature where an agent evaluates a previous experiment, revises its approach based on rubric grades, and repeats until it has completed its objectives. Microsoft-Decision-1 scored as 46 times more consistent than the LLM-based score at three times the speed and resulted in nearly four times the speed on adaptive replanning
Faster, more reliable planning could significantly impact outcomes for long-running scientific experiments.
Those are just a small handful of examples. There are many more potential use cases for Microsoft-Decision-1. Consider trying out the following:
Developers can get started with Microsoft-Decision-1 today in Microsoft Foundry here: aka.ms/decision-1.
Input tokens cost $0.042 USD per million tokens. Output tokens are free.
Now that agentic AI is a reality, we’ve seen that cost plays a major role in how people decide to use AI. And it’s increasingly important to choose the right model for the right job. With agents taking action and making an impact in the real world, decision models have the potential to help people guide and control those agents through complex environments.
We look forward to seeing what developers build with Microsoft-Decision-1 and hearing their feedback. We’ll continue to release updates to the model, including by incorporating evaluations and data to further optimize quality, confidence, and cost.
Quyet-1.0-Large — https://huggingface.co/chinhnc/Quyet-1.0-Large
Surogate Rune 26B-A4B — https://huggingface.co/surogate/rune-26b-a4b-GGUF
GPT-6 Luna Decisions — https://developers.openai.com/api/docs/guides/decisions
deck-31B — https://github.com/krishna-gogineni-765/deck31b
H2O-Lightning-4B — https://huggingface.co/h2oai/h2o-lightning-4b
Strands-Decider 2B — https://github.com/strands-labs/strands-decider