Optimizing the frontier performance curve

@mustafasuleyman
الإنجليزيةقبل 24 ساعة · 29 يوليو 2026
133K
250
32
21
72

ليرة تركية؛ د

Mustafa Suleyman details Microsoft AI's strategy for token efficiency and specialized models, highlighting significant cost savings and the launch of the Frontier Tuning service.

This week MAI-Cyber-1-Flash inside of our MDASH harness, landed at No.1 on the CyberGym benchmark - beating Mythos by 12ppts - at 50% of the cost.

These kinds of performance-cost optimizations are central to the new model deployment paradigm. Tokenmaxxing has been the story of the last few months. Token efficiency is the next big focus across the industry.

To build a frontier firm, you have to optimize frontier performance against cost. Choosing where you want to sit on that curve is critical. By co-optimizing your models, harnesses, and RLEs you can pick a point on the curve that suits your firm.

As we race towards a world where frontier firms deploy always-on, real-time agents operating continuously in the background, inference demand costs are going to skyrocket.

In most cases, frontier generalist models aren’t necessary for every task. By tuning models for a specific product, you can maintain or even exceed frontier performance, while reducing token costs dramatically, sometimes by as much as 50% or more.

As Satya just mentioned on our Q4 Earnings call, we’ve shipped over a dozen specialist MAI models across Microsoft products since last quarter. All of them maintain or improve quality vs. existing deployed models, whilst using significantly fewer tokens, in many cases saving as much as 50-90% of GPU costs.

This is the hill-climbing machine in action. Co-optimizing your own models, in your own harness, with your data and workflows, trained in your RLEs. We believe this is the new rhythm of a frontier firm, and we’re building a service called Frontier Tuning, which will help all firms to work like this.

This flywheel is what will enable the frontier ecosystem to prosper, and be more resilient, ensuring you can swap models, build your own IP, and control costs.

Here’s a list of some of our models that shipped this quarter and token efficiency impact:

  • MAI-Code-1-Flash delivers 10% higher code accept rate than GPT-5.4 Mini and Claude Haiku 4.5 in VS Code, 10% lower median token usage, and increases retention by 9% on average.
  • MAI-Code-1-Flash is also the basis for our Excel model which deliver comparable performance to similar models, whilst being served on H100s, vs. GBs.
  • MAI-Image-2.5-Flash reduces GPU costs by up to 84% in PowerPoint against GPT-Image-2. In OneDrive it delivers up to 2.5x token efficiency while cutting P95 latency by 25% against GPT-Image-1.5, and increasing user save rates by 25%.
  • MAI-Voice-2-Flash delivers world class quality while reducing GPU costs up to 89%, while being 2x faster and 32% cheaper than its predecessor.
  • MAI-Transcribe-1.5 which can run up to 5x faster than competitor models, is shipped to 58 languages across the Dragon Copilot product, which is used by 170k Medical providers, for nearly 30m patient encounters.

It’s been a summer of hard work during the most exciting time in the history of our industry. Awesome work by the team! Much more to come, we’re just getting started.

ريمكس في YouMind

قم بتحويل مقال سريع الانتشار إلى سير عمل كامل المحتوى

قم بتجميع المصدر وفك تشفير النمط وإنشاء الأصول وصياغة القصة وتوزيعها من مساحة عمل واحدة تعمل بالذكاء الاصطناعي.

اكتشف YouMind
للمبدعين

حول Markdown إلى مقالة 𝕏 نظيفة

عندما تنشر كتاباتك الطويلة، فإن الصور والجداول وكتل التعليمات البرمجية تجعل تنسيق 𝕏 مؤلمًا. YouMind يحول مسودة Markdown كاملة إلى مقالة نظيفة وجاهزة للنشر 𝕏.

حاول Markdown إلى 𝕏

المزيد من الأنماط لفك التشفير

المقالات الفيروسية الأخيرة

استكشاف المزيد من المقالات الفيروسية