For most of the generative AI boom, Microsoft's pitch to enterprise customers rested on a simple promise: the smartest models in the world, wired directly into the apps people already used. That promise is now quietly changing. Microsoft MAI models, the company's own in-house systems, have started answering real Copilot prompts inside Excel and Outlook, and the executives running the effort have been unusually blunt about why. The goal is money. Paying OpenAI and Anthropic for every keystroke of assistance has become one of the largest line items in the AI economy, and Microsoft would rather keep that cash inside the building.
- Microsoft has begun routing tens of thousands of weekly Copilot requests in Excel and Outlook through its own MAI models instead of OpenAI or Anthropic.
- AI chief Mustafa Suleyman said the plan is to reduce and eventually eliminate the cost of paying Anthropic.
- Benchmarks put the new reasoning model, MAI-Thinking 1, well behind OpenAI and Anthropic and roughly level with DeepSeek V3.2.
- Microsoft is one of several large firms, alongside Amazon, Uber, Meta, and Accenture, pulling back from aggressive AI spending.
The shift was first reported by Bloomberg and picked up by TechCrunch and The Decoder in early July 2026. According to those reports, Microsoft has started using MAI models to answer a share of user prompts in its productivity apps. The volume is still small relative to the total, but it is no longer a lab experiment. Real customers, doing real work in spreadsheets and inboxes, are already getting answers from a Microsoft model rather than a partner's.
What Microsoft is actually swapping
The change is targeted rather than wholesale. Microsoft is not ripping OpenAI and Anthropic out of Office 365 overnight. Instead it is redirecting the high-volume, low-drama requests, the kind that dominate day-to-day use of Excel and Outlook, toward its cheaper homegrown systems. Summarize this thread. Clean up this column. Draft a short reply. These are the workloads that pile up by the millions and quietly run the meter.
The Decoder reports that Microsoft's own models now handle tens of thousands of requests per week across Excel and Outlook. That figure sounds large until you set it against the scale of Microsoft 365, which serves hundreds of millions of paid seats. The company is testing the waters, watching quality metrics, and expanding the footprint where the results hold up. A separate proprietary transcription model is slated to arrive in Teams, taking over another commodity task that used to lean on outside providers. Meeting transcription is a natural target, because it runs constantly across enterprise accounts and rarely needs frontier-grade reasoning to do its job well.
What makes the move notable is how it reverses Microsoft's earlier marketing. The company had openly advertised that large parts of Office 365 ran on models from both OpenAI and Anthropic, treating that partnership as a selling point. Anthropic's Claude, in particular, had been positioned inside Copilot as a premium reasoning option. Now the same apps are being reworked so that the default answer comes from Redmond's own stack, with the marquee partners reserved for the harder cases.
The Microsoft MAI models lineup
The systems doing this work trace back to Microsoft's Build conference in June 2026, where the company introduced a family of seven new MAI models. CNBC covered the launch, framing it explicitly as a way to lessen reliance on OpenAI and lower costs for developers. The lineup spanned general-purpose text models plus an agentic coder and a text-to-image generator, giving Microsoft a slate broad enough to cover most of what a productivity suite needs without phoning a partner.
The headline release was MAI-Thinking 1, Microsoft's first reasoning model, built to work through multi-step problems rather than answer in a single pass. Microsoft says it trained the model on roughly 30 trillion tokens of commercially licensed data and did not distill it from GPT or Claude outputs, a detail the company clearly wants on the record given the licensing fights swirling around training data. In practice that means MAI-Thinking 1 is meant to stand on its own rather than borrow its intelligence from the very rivals Microsoft is trying to stop paying.
Microsoft has also pointed to real tuning work behind the models. In one example cited in the reporting, the company refined its systems around the needs of consulting firm McKinsey and claimed it could match or beat a comparable third-party option at a fraction of the running cost. That is the whole thesis in miniature. If a fine-tuned in-house model can hit an acceptable quality bar on a narrow task for a tenth of the price, the economics of routing that task to an outside lab stop making sense.
Suleyman's cost math
Mustafa Suleyman, the former DeepMind co-founder who now runs Microsoft AI, has not hidden the motive. In comments relayed by The Decoder, he said the quiet part plainly: "We pay a lot of money to Anthropic, so our goal is to reduce and ultimately eliminate that cost." The phrasing matters. This is not a hedging strategy or a bit of vertical integration around the edges. It is a stated intent to remove a supplier from the bill entirely, on the workloads where Microsoft can afford to.
The numbers underneath that ambition are the reason it exists. Frontier model access is priced per token, and the tokens add up fast when a feature is embedded in software that hundreds of millions of people touch every day. To put the input side in perspective, a top-tier frontier model can cost several dollars per million input tokens and many times that on output. Every autocomplete, every summary, every rewrite draws from that pool. Multiply it across an install base the size of Office and the annual figure becomes the kind of number a CFO circles in red.
Suleyman's framing also reflects a strategic worry that goes beyond any single invoice. For years Microsoft's AI story was inseparable from OpenAI, its closest partner and largest model supplier. Leaning on a rival like Anthropic for Copilot's reasoning added a second dependency. Owning the models outright, even middling ones, gives Microsoft a lever it did not have before: the ability to walk away from a price it does not like.
Where the MAI models still trail
The catch is quality, and Microsoft has not pretended otherwise. Benchmarks referenced by The Decoder place MAI-Thinking 1 behind the reasoning models from OpenAI and Anthropic by a wide margin, landing roughly on par with DeepSeek V3.2, the open-weight Chinese model that has become a common yardstick for capable-but-cheap systems. That is a revealing comparison. DeepSeek V3.2 is a respectable model, but no one confuses it with the current frontier, and Microsoft is effectively telling customers that its default Copilot brain now sits in that tier for the tasks being switched over.
It is worth being precise about what parity with DeepSeek V3.2 implies. That model earned attention for delivering strong results on coding and reasoning tests at a running cost far below the American frontier, which is exactly why it works as a reference point for cheap capability. Sitting alongside it means MAI-Thinking 1 is genuinely useful, capable of structured multi-step answers, yet clearly a rung below the systems that top the public leaderboards. Microsoft is choosing that rung deliberately for the traffic it moves first, because the marginal quality it gives up is small next to the cost it saves, and because the tasks in question rarely stress a model to its limit.
For a large slice of everyday work, that gap may not bite. Summarizing an email chain or tidying a table does not demand the strongest model ever built, and users often cannot tell which system answered. The risk sits at the edges, on the prompts where reasoning depth actually decides whether the output is useful or subtly wrong. The Decoder put the customer-facing tradeoff bluntly: people could end up with weaker AI while paying the same subscription price they always did. Microsoft is betting that most users will not notice, or will not mind, on the workloads it moves first.
Microsoft declined to elaborate. Asked about the changes, the company told TechCrunch it had nothing further to share, and it has not published a public breakdown of which prompts route to MAI versus a partner model.
An industry that stopped tokenmaxxing
Microsoft is not moving in isolation. TechCrunch framed the shift as the company joining a broader retreat from the free-spending posture that defined the previous two years, when the reflex was to throw the biggest available model at every problem and worry about the bill later. That reflex even earned a nickname in some circles: tokenmaxxing, the habit of maximizing model calls and context length with little regard for cost.
The reporting names a cluster of large firms tightening up at the same time. Amazon and Uber have both signaled more disciplined AI budgets. Accenture has reined in its own spending, and Meta has been reworking how aggressively it pays for inference. The common thread is that the experimental phase is ending. Companies have shipped enough AI features to know which ones people actually use, and they are now asking a colder question about each one: does the value it delivers justify what it costs to run. When the answer is no, the workload either gets cut or gets moved to a cheaper model.
That logic favors exactly the kind of vertical integration Microsoft is pursuing. Cloud providers and platform owners have a structural advantage, because they can build their own models and route the commodity traffic internally while keeping premium options available for anyone willing to pay more. Smaller companies without that option face a starker choice between absorbing the cost and degrading the feature. The pattern also helps explain the wider surge of interest in efficient open-weight models, which offer a way to escape per-token pricing without building a lab from scratch.
The bill on the other side
Every dollar Microsoft stops paying is a dollar that stops flowing to a model lab, and the loss is not evenly spread. Microsoft has long been OpenAI's biggest backer and a heavy consumer of its models, but the Anthropic spend is newer and more exposed. Claude found its way into Copilot as a premium reasoning option precisely because enterprise customers wanted variety and quality, and that placement turned Microsoft into a meaningful revenue channel for Anthropic almost overnight. Suleyman naming Anthropic by name, rather than speaking about outside models in the abstract, reads as a signal about which relationship is most in play.
The exposure is uncomfortable because the labs need this kind of embedded, high-volume distribution to justify their valuations. Frontier training runs cost billions, and the way that spend gets recouped is by placing models inside software people use constantly. When the largest software distributor on earth starts building substitutes for the routine end of that traffic, it removes the steadiest and most predictable part of a lab's revenue while leaving the harder, spikier demand behind. A lab can survive losing the easy prompts, but it cannot pretend the loss is nothing.
There is a strategic irony worth sitting with. The same open-weight models that made Microsoft's move possible, systems trained by others and released under permissive licenses, are the ones eroding the pricing power of the frontier labs from below. Microsoft did not have to match OpenAI or Anthropic to justify the swap. It only had to reach the level of a strong open model like DeepSeek V3.2, then apply its own tuning and distribution muscle on top. That lowers the bar for anyone else considering the same path, and it suggests the pressure on premium per-token pricing is structural rather than a one-off negotiating tactic. For readers tracking how these economics ripple outward, our ongoing coverage follows the same theme across the open-weight ecosystem.
What Copilot buyers may notice
The most concrete change for customers may be how they are billed, not just which model answers. CEO Satya Nadella has hinted at a move toward usage-based pricing, where the cheaper MAI models serve as the default and access to a top-tier OpenAI or Anthropic model becomes a premium add-on that costs extra. In that world the flat Copilot subscription buys you Microsoft's own intelligence, and the frontier becomes an upsell.
This is a familiar arc in software, where the shift from flat rates to metered usage-based pricing tends to follow the moment a vendor's own costs become variable and hard to predict. For buyers, it introduces a new calculation. The question stops being whether Copilot is worth a fixed monthly fee and becomes which tier of intelligence each task deserves, and whether the default tier is good enough for the work at hand. Power users who lean on Copilot for genuinely hard reasoning may find themselves paying more to keep the quality they had, while lighter users pay the same or less for a service that quietly changed underneath them.
There is a trust dimension too. Microsoft spent two years telling enterprise buyers that Copilot ran on the best models available, and some of those buyers made purchasing decisions on that basis. Swapping in a weaker default without a loud announcement, however sound the cost logic, invites the question of what else might change silently inside a product people rely on daily. Microsoft's decision to say little when asked does not help on that front. Coverage from SiliconANGLE and others has already framed the move as much about supplier leverage as about customer benefit.
A test of how good is good enough
The deeper story here is not really about Microsoft's spreadsheet features. It is about a question the whole industry is now forced to answer with real money on the line: how good does a model need to be for a given job. For two years the assumed answer was as good as possible, because capability was the product and cost was a problem for later. Microsoft's swap is a wager that later has arrived, and that for the bulk of everyday tasks a merely competent model at a tenth of the price wins.
If that wager pays off, it reshapes the economics for everyone selling frontier access. The most lucrative slice of the market, the endless stream of routine prompts embedded in mass-market software, starts draining toward cheaper in-house and open-weight systems, leaving the top labs to compete over the harder and rarer work where their edge still shows. If it fails, and customers revolt over degraded output or confusing new pricing, Microsoft will have taught its rivals exactly which corners not to cut. Either way, the era when the answer to every AI question was to buy the biggest model available is closing, and Microsoft, of all companies, is the one holding the scissors. The rest of the sector will be watching Excel and Outlook to see how the experiment goes.