
Microsoft-Decision-1 Is Live on Foundry. Its Benchmarks Tab Is Empty.
- OrcaNEWOrca: OrcaCyber Zero 1.52026-10-10$3.00 / $7.50 per 1M tokens · 90 tok/s
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 121 tok/s
- OpenAIOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAIOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAIGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 53 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 423 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 60 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 361 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 231 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
The most interesting thing about this week's Microsoft release is not what Microsoft-Decision-1 can do. It is what Microsoft chose not to print about it. The model is live: it went generally available on Microsoft Foundry on October 8, 2026, two days before this was written. It is a decision-scoring model — you give it a state and a question with a fixed set of answers, and it returns a calibrated probability per answer — post-trained by Microsoft on the open-weight Qwen3.5-9B, running one pass over up to 32,768 tokens and emitting zero output tokens because it never generates anything at all. The catalogue page has a Benchmarks tab. It contains a methodology paragraph and no figures.
That gap is the story, and it is a more useful one than another "Microsoft ships a model" item. Every serious question about a scorer is a calibration question — does a returned 0.8 mean 0.8 — and a launch that arrives without a single Brier score or expected-calibration-error figure leaves the one number that matters to be measured by whoever adopts it. What follows is what the Foundry page actually documents, what it conspicuously leaves out, and what an evaluation team can do about it this week.
What shipped, precisely
Microsoft-Decision-1 is a hosted API. The contract is one call in, one distribution out, with no decoding loop anywhere in the path: the request carries the material to be judged plus a question with a bounded answer set, and the response carries a probability for each option. Microsoft lists the supported question shapes as yes/no, multiple-choice, rating, classification and rubric-based, all within a single invocation of up to 32K tokens. It is text-only — no image, audio or video input, and nothing but numbers out.
The exclusions are stated as plainly as the features, and they are worth reading before anything else: not designed for text generation, open-ended question answering, conversation, translation or summarization, and not intended for tasks that require knowledge absent from the input. It produces no rationales. The published use cases are all places a platform team already has a labelled decision to make — grading a generated answer against a rubric, judging retrieval relevance, triaging a queue, gating a proposed agent tool call, screening content against thresholds the application defines rather than a fixed vendor policy, and auto-accepting high-confidence outcomes while escalating the rest.
Two operational details stand out from the deployment listing. The first is that Microsoft explicitly supports an abstention option such as "cannot tell" when the supplied evidence is insufficient — that is the difference between a scorer that is calibrated and one that is merely confident, and it is what makes thresholding work. The second is that batch inference is disabled. You cannot amortize a large scoring run through the batch channel the way you would with a generative model, so per-call latency is the latency of your pipeline rather than an offline job's problem.
Distribution is Foundry-only, in the "Direct from Azure" portfolio as a serverless or unified-endpoint deployment on standard SKU — pay-as-you-go or reserved provisioned throughput. Weights are not distributed. There is no Hugging Face repository, no download, no fine-tuning path, and no self-hosting option. Applications integrate over HTTPS with standard Azure authentication. The training disclosure reports the dataset was first used in September 2026 with collection ongoing, which is the shortest possible distance between training data and a GA date and is normal for a post-train on someone else's released base.
The benchmarks tab, quoted in full
Here is the entirety of what Microsoft has published about how well the model performs. The evaluation used "public and community decision benchmarks and held-out internal test sets not used in training." The metrics were accuracy, calibration error, safety recall, false-positive rates and fairness consistency. Option order was varied. Paired statistical tests were applied. The claim is qualitative: Microsoft-Decision-1 "performs on par with leading decision models and ahead of other open decision models evaluated with the same methodology."

That is a competent evaluation design described without a result. It is not an accusation to point that out — a methodology paragraph with no table is a specific, checkable choice, and it is a different choice from the one the rest of this small category has made. The open decision models Microsoft is implicitly comparing against publish their numbers: InternLM's Intern-Decision family prints Brier and expected-calibration-error figures on its model cards, TypeSafe's Jev publishes both, and Liquid AI's d1 line ships accuracy tables with its weights. Microsoft is the largest company in this group and the only one asking to be taken on trust.
The company does say where it believes the model is strong and weak, which is more actionable than a headline score. Strongest: reasoning, rule application and robustness to prompt formatting. Competitive: classification, retrieval, fairness, tool use and most multilingual tasks. Weakest: specialized domain knowledge. The self-reported limitations are candid in the same way — scores can shift with phrasing and option ordering, a poorly framed question still returns a score, calibration is strongest on familiar task types, and there are no explanations to audit when an answer looks wrong.
Language coverage carries the same shape of caveat. 25 languages are listed as supported, spanning Japanese, Korean, Arabic, Vietnamese, Thai, Turkish, Hindi, Bengali, Swahili, Hebrew, Persian and Ukrainian among others, with the explicit warning that coverage, quality and calibration "may vary by language" and that non-English, particularly lower-resource languages, is an area of underperformance. The Qwen3.5-9B base supports well over 200 languages. Post-training kept roughly a quarter of that, and the quarter is where the calibration was fitted.

No price on the page either
The catalogue's pricing field does not print a rate. It links out to Microsoft's own model pricing surface, so the per-decision cost is something you read off Azure or off a bill rather than off the model card. For anyone modelling cost per decision at volume that is a real gap, and it is worth stating plainly instead of estimating. Two things do follow from the architecture and are worth carrying into that estimate: 0% of a call's cost is output tokens, because there are none, and the option set is part of the input, so a question with sixty-two descriptive options costs more per call than a yes/no — you are paying for the rubric you wrote, not for the answer.
Why this shape of release is the interesting part
A decision scorer is a bet that the primitive enterprises actually need is not a better writer but a cheaper, more reliable judge. That bet only pays if the probability is trustworthy, because everything downstream of a scorer is a threshold: 0.7 escalates to a person, 0.95 auto-accepts, and the cost of getting that line wrong is paid in bad automated decisions rather than in tokens. A vendor that ships the scorer without the calibration table is asking each customer to re-derive it on their own data.
Microsoft's own documentation recommends exactly that, which softens the criticism and sharpens the practical conclusion at the same time. Validate on data representative of your use case. Set thresholds from the cost of your errors rather than from a default. Always include an abstention option. Randomize option order where ordering could bias the answer. Keep a human in the loop for anything consequential. That is sound advice for any scorer. It is the only advice available for this one.
Trying it this week without committing
The cheapest evaluation is to pick a decision you already make by hand, assemble two hundred labelled cases with the answer sets your application would actually supply, and run them through a Foundry deployment. Compute expected calibration error on the output and you will know more about Microsoft-Decision-1 than anything Microsoft has published about it, because you will have measured it on your distribution rather than on a held-out internal test set. That is an afternoon of work and it retires the entire benchmark question.
Where OrcaRouter fits is on the other half of a scoring loop, and it is the half that generates. We do not host Microsoft-Decision-1 and it is not in our catalogue — a model that returns probabilities instead of text is not something you route chat completions to, and nothing here should be read as an availability claim. What sits behind our one OpenAI-compatible key is the pool of more than 200 models that does the writing: the model that drafts the rubric, the two that produce candidate answers, the one that emits the tool call Decision-1 then scores before it runs. Provider list price is passed through at 0% markup, so a price cut on the generator side is live on ours the same day, and automatic failover keeps the generation leg alive when a single provider degrades — which matters more in a pipeline that scores everything it sees than in one that answers a user occasionally. If you would rather not pick a single judge, the routing DSL composes several models into one call, and model fusion reports their agreement as a scored field rather than as prose you have to read.

What would change this article is a table. Publish the Brier scores and the ECE, or let an independent run land on a leaderboard, and the evaluation above becomes a confirmation instead of the only evidence in existence. Until then the accurate description of Microsoft-Decision-1 is narrow: the weights are real, the contract is documented better than most hosted releases manage, the abstention option is designed in rather than bolted on, and the performance claim is a sentence — a well-written one, attached to no numbers at all.
Bottom line
Microsoft-Decision-1 reached general availability on Microsoft Foundry on October 8, 2026 as a text-only, 32,768-token decision scorer built on Qwen3.5-9B, returning calibrated probabilities over your own option sets with zero output tokens and no weights to download. Its strengths are a clean single-pass contract, a designed-in abstention path, and Azure authentication, billing and governance attached; its weakness is that nobody outside Microsoft has a published number for how well calibrated it is, including Microsoft. Treat the launch as an API becoming available, not as a capability being established, and put your own labelled cases through it before anything downstream depends on a threshold.
