Microsoft has entered the decision-model race — and it did so by skipping its closest AI partner entirely. Microsoft-Decision-1, quietly listed in Foundry and announced by Satya Nadella on October 9, is a small model that doesn’t write text at all: it reads a situation and a fixed set of options and returns calibrated probabilities for each, for use in routing, classification, verification and agent control. Microsoft post-trained it on Alibaba’s open-weights Qwen3.5-9B, three days after OpenAI opened its rival Decisions API to every developer in public beta.
🔍 THE BOTTOM LINE
The significance here is strategic, not technical. Microsoft had every reason to wait for its partner’s own decision-product — instead it built on Alibaba’s open weights, matching category-creator Jev’s $0.042 per million input token price exactly. When the company that owns the distribution layer (Copilot, GitHub, Xbox, Azure) starts routing agent decisions in-house on a cheap Chinese base model rather than OpenAI’s, the OpenAI–Microsoft relationship has quietly become more competitor than partnership.
What a “decision model” actually is
A decision model answers typed questions about a given state with a yes/no probability, a choice from options you define, or a position on an ordered rubric — through a structured Decisions API rather than a chat endpoint. Your code owns the workflow and acts on the answers: act, defer, retry, escalate, or hand the task to a person. Microsoft’s model card lists grading of AI responses, grounding checks against supplied evidence, explicit “cannot tell” abstentions, agent guardrails, data labelling, safety screening, and even robotics among the target uses. Output is JSON; weights update continually while the API shape stays fixed.
The category is barely a month old. As we covered when Amazon open-sourced its 2B Decider, Jev kicked things off in mid-September with a model that judges options instead of writing text, and within weeks OpenAI, Cloudflare (Clef), Perplexity, Upstage, AWS and now Microsoft had all shipped their own. TypeSafe, Jev’s creator, says nearly 30% of the Fortune 500 has tried Jev in three weeks — and on October 9, the same day Microsoft launched Decision-1, TypeSafe closed an $870 million Series A at a $7.5 billion valuation.
What Microsoft claims, and where the catch sits
Microsoft says Decision-1 posted the highest accuracy in a 36-benchmark, roughly 150,000-question evaluation it ran itself — 83.5% average, ahead of Quyet-1.0-Large at 81.9% — at 85ms p50 latency, 4.5 times faster than the runner-up and 35 times faster than GPT-6 Sol. Nadella’s post says the company is already testing it across Microsoft. The internal case studies are the real product pitch: Xbox Research used it to sort more than 10,000 pieces of player feedback at 14 times the speed and 200 times less cost than GPT-6 Sol, the Copilot team found it competitive with GPT-5.6 Luna at 100 times the speed, and Microsoft Discovery scored agent experiments 46 times more consistently than LLM-based scoring.
Every efficiency number is vendor-run, and rivals were not all in the comparison set — Cloudflare’s open-source Clef models, built on the same Qwen family, weren’t included. As The Decoder summarises, the Foundry listing is text-only up to 32,768 tokens (OpenAI’s API and Cloudflare’s Clef both handle images), no open weights have been announced while Clef ships under Apache 2.0, and Microsoft hasn’t confirmed compatibility with the System One API that AWS, Upstage and Ollama have already adopted. Decision-1 is live on both Foundry and OpenRouter at $0.042 per million input tokens, with output free.
The partner it chose not to use
The layer beneath the benchmarks is where this story actually gets interesting. Microsoft post-trained Decision-1 on Alibaba’s Qwen3.5-9B — the same base Cloudflare picked for Clef-flash and the same base Cognition ported its open Kev models to for roughly $95 of rented GPU time. The economics are the point: at Qwen3.5-9B’s quality level, a competent lab can build a competitive decision model in weeks for almost nothing, which is why Jev’s price instantly became the going rate rather than a moat. Microsoft says it plans to rebase future versions on its own MAI models and OpenAI’s — the Qwen choice reads as a speed decision, not a partnership verdict.
But it happened in the same week OpenAI’s own decision product went public beta, and Microsoft shipped past it. There’s an unresolved thread here: on Wednesday, Microsoft said GitHub Copilot will soon decide when to run a task on-device versus sending it to cloud models — model routing is a listed Decision-1 use case, and Microsoft hasn’t said whether Decision-1 will make those calls. With routing decisions to make across Copilot, GitHub and Xbox, Microsoft has every incentive to keep them in-house. For a partner relationship, read that carefully.
The calibration question
The open risk in this whole category is that calibrated scores go soft under adversarial pressure. In the JevOut preprint (arXiv:2609.30243), USC researcher Zixiang Xu and co-authors showed that short, natural-sounding additions to the context flipped 312 of Jev’s 508 initially correct decisions — often at high confidence. Microsoft says Decision-1’s decisions flipped on just 1.3% of eight perturbation tests — with zero flips on paraphrasing, reversal or option shuffling — but that’s Microsoft’s own testing, and its Foundry documentation advises customers to validate calibration on their own data. Independent JevBench results are pending.
The pattern will be familiar to anyone who read our analysis of agents cheating their own benchmarks: whenever a model’s score becomes the signal a system acts on, someone writes an input designed to move the score. Decision models put that signal directly in the automation path — one probability gate standing between an agent and “just do it.”
❓ FAQ
What is Microsoft-Decision-1 for? Routing, classification, prioritisation, verification, workflow control, agent guardrails, data labelling and AI judging. It answers structured questions with calibrated probabilities — your application decides what to do with the answer. It is not built for open-ended generation or conversation.
Why is building it on Alibaba’s Qwen notable? Qwen3.5-9B is an open-weights model anyone can license and post-train, and it’s the same base Cloudflare’s Clef-flash used. The choice let Microsoft ship in days and match Jev’s price — but it means Microsoft’s newest agent-control layer isn’t built on OpenAI technology, days after OpenAI’s competing Decisions API launched.
How much does it cost and where is it available? $0.042 per million input tokens with free output tokens — identical to Jev’s price — through Microsoft Foundry and OpenRouter.
Can Microsoft’s accuracy claims be trusted? Treat them as marketing-adjacent until independent results land. The 36-benchmark evaluation was run by Microsoft itself, Cloudflare’s Clef was excluded, and Microsoft’s own documentation advises validating calibration on your own data. Independent JevBench results are pending.
🔍 THE BOTTOM LINE
Decision models are becoming the plumbing of the agent economy — the cheap, fast gate that decides whether an AI acts — and Microsoft just grabbed a seat at the table by building on someone else’s open weights rather than waiting for its partner. The open question this week isn’t whether Decision-1 is good; it’s what it means that the world’s biggest cloud vendor and OpenAI’s closest partner chose Alibaba’s Qwen over its own ally for a product category OpenAI just entered. Watch whether Decision-1 ends up routing Copilot’s traffic. If it does, the partnership has a new, quieter shape.
📰 Sources
- Microsoft Foundry — Microsoft-Decision-1 model card (October 2026)
- Satya Nadella — announcement post on X (October 9, 2026)
- The Decoder — Microsoft’s Decision-1 model enters the fast-growing AI decision model race (October 10, 2026)
- The New Stack — Microsoft skipped OpenAI’s decision model and built its own on Alibaba’s Qwen (October 10, 2026)
- TestingCatalog — Microsoft launches Decision-1 model in Foundry (October 2026)
- MarkTechPost — Microsoft AI Releases Microsoft-Decision-1 (October 9, 2026)