Microsoft launches new in-house AI models it says cut costs up to 89% versus OpenAI

Read full story on VentureBeat
Share
Microsoft launches new in-house AI models it says cut costs up to 89% versus OpenAI
AI disclosure

Summary

<p><a href="https://microsoft.ai/">Microsoft AI</a> released two new in-house models into public preview on Wednesday — <a href="https://microsoft.ai/news/introducing-mai-image-2-5-pro-and-mai-voice-2-flash/">MAI-Image-2.5-Pro</a>, its highest-fidelity image generator to date, and <a href="https://microsoft.ai/news/introducing-mai-image-2-5-pro-and-mai-voice-2-flash/">MAI-Voice-2-Flash</a>, a speech model built for high-volume enterprise workloads — while publishing production data that amounts to the company&#x27;s most aggressive argument yet that it can power its own products without leaning on OpenAI&#x27;s frontier models.</p><p>The announcement, made by <a href="https://microsoft.ai/">Microsoft AI&#x27;s Superintelligence team</a>, lands roughly a year after the company committed to building purpose-built models internally, and it arrives with an unusual level of specificity about where those models now run: <a href="https://www.bing.com/">Bing</a>, <a href="https://www.microsoft.com/en-us/microsoft-365/powerpoint">PowerPoint</a>, <a href="https://www.microsoft.com/en-us/microsoft-365/onedrive/online-cloud-storage">OneDrive</a>, <a href="https://www.microsoft.com/en-us/dynamics-365">Dynamics 365</a>, <a href="https://excel.cloud.microsoft/en-us/">Excel</a>, <a href="https://github.com/features/copilot">GitHub Copilot</a>, and <a href="https://azure.microsoft.com/en-us">Azure</a>. The message to enterprise buyers — and, implicitly, to OpenAI — is that Microsoft&#x27;s homegrown models are no longer research projects. They are production infrastructure serving millions of users.</p><p>&quot;Each of these enhancements is a step toward the same goal: Microsoft products, powered by Microsoft models,&quot; the company wrote in its announcement blog.</p><h2><b>How MAI-Image-2.5-Pro and MAI-Voice-2-Flash stake out opposite ends of the AI cost curve</b></h2><p>The two new releases occupy opposite ends of what Microsoft calls the quality-speed-cost curve, and the positioning is deliberate. <a href="https://microsoft.ai/news/introducing-mai-image-2-5-pro-and-mai-voice-2-flash/">MAI-Image-2.5-Pro</a> targets the premium tier: hero imagery, detailed editing, and precise in-image text rendering — the last of which has long been a notorious weak spot for image generation models. Microsoft priced the model at $5 per million text input tokens, $8 per million image input tokens, and $106 per million image output tokens. The base <a href="https://microsoft.ai/news/introducing-mai-image-2-5-pro-and-mai-voice-2-flash/">MAI-Image-2.5</a> model recently launched at <a href="https://microsoft.ai/news/introducing-mai-image-2-5/">No. 2 for image editing on Arena</a>, the community leaderboard that has become a de facto scoreboard for generative media.</p><p>The creative industry appears to be taking notice. Rob Reilly, global chief creative officer at advertising giant WPP, called the Pro model &quot;a strong leap forward for GenMedia tools&quot; in a statement included in Microsoft&#x27;s announcement, adding that &quot;Microsoft has firmly established itself among the leaders in generative AI.&quot;</p><p><a href="https://microsoft.ai/news/introducing-mai-image-2-5-pro-and-mai-voice-2-flash/">MAI-Voice-2-Flash</a> goes the other direction. First previewed at Microsoft&#x27;s <a href="https://news.microsoft.com/build-2026/">Build conference</a>, Flash runs twice as fast as MAI-Voice-2 and costs 32% less, priced at $15 per million characters. It is designed for the unglamorous but enormous market of high-volume voice — call centers, voice agents, and real-time speech applications where latency and cost-per-call matter more than marginal gains in expressiveness. Together, the two models reflect a strategy of building families of models rather than a single flagship, because, as the company put it, a creative studio chasing maximum fidelity has very different needs from a customer service operation handling millions of calls a day.</p><h2><b>Microsoft&#x27;s production metrics show in-house models cutting GPU costs by up to 89%</b></h2><p>The model launches are arguably less newsworthy than the deployment metrics Microsoft attached to them — numbers that read like a systematic case for swapping out third-party frontier models across its product portfolio. </p><p><a href="https://explore.microsoft.com/en-us/bing/features/bing-image-creator?form=MA13FV">Bing Image Creator </a>now runs entirely on <a href="https://microsoft.ai/news/introducing-mai-image-2-5-pro-and-mai-voice-2-flash/">MAI-Image-2.5</a>, end to end, marking the first time the consumer image tool is fully in-house. In PowerPoint, Microsoft says MAI-Image-2.5 reduces GPU costs by up to 84% compared with GPT-Image-2, OpenAI&#x27;s image model. In OneDrive, where MAI-Image-2.5 is now the default for key image-editing scenarios, the company reports a 26% increase in save rates, roughly 25% lower P95 latency, and 2.5 times greater efficiency under medium-utilization production workloads.</p><p>On the voice side, <a href="https://microsoft.ai/news/introducing-mai-image-2-5-pro-and-mai-voice-2-flash/">MAI-Voice-2-Flash</a> now powers Dynamics 365 Contact Center — the platform used by customers including T-Mobile and EasyJet — where Microsoft claims GPU cost reductions of up to 89%. The model is also integrated into Azure Voice Live for developers building speech-to-speech agents.</p><p>Perhaps the most consequential deployment sits in healthcare. Microsoft&#x27;s <a href="https://www.microsoft.com/en-us/health-solutions/clinical-workflow/dragon-copilot">Dragon Copilot</a>, used by 170,000 medical providers and responsible for processing 28 million patient encounters last quarter, now runs on MAI-Transcribe-1.5 for its multilingual workflow across 58 languages. Microsoft says internal evaluations show a 50% relative reduction in both transcription and language-identification error rates across most languages — a meaningful claim in a domain where transcription errors can propagate directly into clinical notes.</p><h2><b>Inside the &#x27;hill-climbing&#x27; strategy that lets small models beat GPT-5.6 in Excel</b></h2><p>In a companion post published the same day, Microsoft detailed the methodology behind these results — what it calls its &quot;<a href="https://microsoft.ai/news/hill-climbing-mai-models-for-github-copilot-and-excel/">hill-climbing machine</a>,&quot; an integrated flywheel of data, models, and the product &quot;harness&quot; that surrounds them.</p><p>The clearest example is <a href="https://microsoft.ai/news/introducingmai-code-1-flash/">MAI-Code-1-Flash</a>, the lightweight coding model launched in GitHub Copilot in June. Microsoft says the model achieves an approximately 10% higher code accept rate than GPT-5.4 Mini and Claude Haiku 4.5 in VS Code, while using 10% fewer median tokens. Developer retention tells a similar story: users were 6% more likely to return across multiple days than with GPT-5.4 Mini, and 11% more likely than with Claude Haiku 4.5.</p><p>Then Microsoft did something more interesting. It took the MAI-Code-1-Flash checkpoint and further <a href="https://microsoft.ai/news/hill-climbing-mai-models-for-github-copilot-and-excel/">trained it inside an Excel reinforcement learning environment</a>, teaching a coding model the tools and workflows of spreadsheet knowledge work. The result, according to production user feedback, is a model on par with GPT-5.6 for the most common Excel tasks — while being small enough to run on Nvidia&#x27;s older H100 and even A100 GPUs rather than requiring the latest-generation accelerators.</p><p>That hardware detail deserves emphasis. Every major AI company is fighting for allocation of cutting-edge chips, and a model that delivers frontier-adjacent quality on two-generation-old silicon fundamentally changes the deployment economics. It also frees the newest hardware — including Microsoft&#x27;s now-operational GB200 cluster — for training rather than serving.</p><h2><b>Satya Nadella&#x27;s &#x27;frontier diffusion&#x27; manifesto redraws the OpenAI relationship</b></h2><p>Microsoft CEO Satya Nadella framed the announcements in a lengthy post on X titled &quot;<a href="https://x.com/satyanadella/status/2080329851127669104">Frontier Diffusion &amp; Control</a>,&quot; which functions as something close to a strategic manifesto. &quot;We can now take saturated frontier capabilities and deliver them at scale and at lower cost through models optimized for high-usage products, while continuing to use frontier models for frontier needs,&quot; Nadella wrote, adding that Microsoft is &quot;beginning to route traffic across our first-party surfaces to MAI whenever our models match or outperform frontier alternatives.&quot;</p><p>Translated from executive prose: capabilities that were state-of-the-art a year ago are now table stakes, and Microsoft believes it can replicate them cheaply for the specific, repetitive tasks that dominate real product usage. Why pay frontier prices for a frontier model when a user just wants to reformat a spreadsheet column?</p><p>Nadella was careful to note that &quot;frontier models from OpenAI and Anthropic are part of the orchestration system alongside MAI&quot; — but he also articulated a pointed principle of model independence, arguing that a company&#x27;s evaluations &quot;should continue to hill climb even when any given model has been removed.&quot; </p><p>“Keeping the harness, memory, context, and skills outside the model, he argued, is what gives Microsoft control. The subtext is hard to miss. Reuters reported in April that Microsoft’s <a href="https://www.reuters.com/legal/litigation/microsoft-end-exclusive-license-openais-technology-2026-04-27/">exclusive license to OpenAI’s technology</a> had been revised into a non-exclusive arrangement, and The Information reported last September that Microsoft had <a href="https://www.theinformation.com/articles/microsoft-buy-ai-anthropic-shift-openai">begun incorporating Anthropic models</a> into some products. Wednesday’s announcement completes the triangle: Microsoft as orchestrator, with its partners’ frontier models as interchangeable components and its own models absorbing an ever-larger share of routine traffic.”</p><h2><b>Developers cheer cheaper task-specific models while skeptics question Microsoft&#x27;s track record</b></h2><p>The response online captured both the appeal and the skepticism surrounding the strategy. &quot;I love when people use small models for niche tasks,&quot; wrote one X user, <a href="https://x.com/mavihsk/status/2080330529547993252">@mavihsk</a>, responding to Nadella&#x27;s post. &quot;Why do I have to use the all-knowing model just to change my field in Excel?&quot; Another user, <a href="https://x.com/nabu_lines/status/2080343512780837226">@nabu_lines</a>, distilled the pitch neatly: &quot;cost and performance both improve when you stop overusing the biggest model.&quot;</p><p>Others were less charitable about Microsoft&#x27;s execution track record. &quot;Microsoft is the worst when it comes to listening to user feedback,&quot; wrote designer <a href="https://x.com/designedbyabin/status/2080332368301412434">@designedbyabin</a>, arguing the company &quot;will lose the AI race because they repeatedly failed to understand user needs.&quot; And one user, <a href="https://x.com/tokenoverflow/status/2080386145712824694">@tokenoverflow</a>, offered a drier critique of the model-independence pitch: &quot;i want it keep hill climbing after removing microsoft.&quot;</p><p>The skeptics raise a fair point. Microsoft&#x27;s self-reported metrics — accept rates, save rates, GPU savings — come from its own internal evaluations, not independent benchmarks, and the company chooses which comparisons to publish.</p><p>But the strategy&#x27;s logic does not depend on any single number. Nadella&#x27;s framing that software now has &quot;<a href="https://x.com/satyanadella/status/2080329851127669104">real marginal cost for the first time</a>&quot; explains why Microsoft is obsessive about tokens, GPUs, and serving costs: when AI features run on every keystroke across a billion-user product portfolio, an 84% GPU cost reduction is not an optimization. It is the difference between a viable business and a money pit.</p><h2><b>Why Microsoft is turning its internal AI playbook into an Azure product</b></h2><p>The final piece of the strategy is that Microsoft is selling the playbook, not just the models. Nadella explicitly positioned the hill-climbing approach as &quot;a template for every other AI native, SaaS, or Enterprise company,&quot; and Microsoft is packaging the toolchain through Foundry and what it calls Frontier Tuning — letting enterprises train specialized models against their own proprietary evaluations and reinforcement learning environments. That turns Microsoft&#x27;s internal cost-cutting exercise into an Azure product, and it gives enterprise customers a reason to run their AI workloads on Microsoft&#x27;s cloud even if the models themselves come from elsewhere.</p><p>The company&#x27;s emphasis on models trained &quot;on clean, traceable, enterprise-grade data, without distillation from third-party models&quot; serves the same commercial end. In an industry facing mounting scrutiny over training data provenance, Microsoft is betting that enterprise buyers — and courts — will care where model capabilities come from. Microsoft says it is now extending the hill-climbing approach to <a href="https://copilot.microsoft.com/">Copilot Chat</a>, <a href="https://outlook.live.com/mail/">Outlook</a>, and <a href="https://www.microsoft.com/en-us/microsoft-365/powerpoint">PowerPoint</a>, and both new models are available in public preview through <a href="https://azure.microsoft.com/en-us/products/ai-foundry">Microsoft Foundry</a> and the <a href="https://playground.microsoft.ai/">MAI Playground</a>. &quot;None of this is an endpoint,&quot; the company wrote. &quot;We&#x27;re just getting started.&quot;</p><p>Seven years ago, <a href="https://www.cnbc.com/2024/08/10/rise-of-openai-microsofts-13-billion-artificial-intelligence-bet.html">Microsoft bet more than $13 billion</a> that OpenAI would build the future of AI. Wednesday&#x27;s announcement suggests the company has since learned a cheaper lesson: the future of AI may belong to whoever builds the frontier, but the profits belong to whoever makes it ordinary.</p>

Discussion on

Trending posts from X.

Original reporting

Open original source

Related coverage

Read full article on VentureBeat

Get the AFBytes Brief

Major stories, AI-assisted analysis, and what to watch next. Free, monthly, unsubscribe anytime.