跳到正文
原文
Microsoft AI:官方博客(网页)·· 2026-07-24精选AI 评分63

Microsoft 用 hill-climbing 方法推出面向 GitHub Copilot 和 Excel 的专用 MAI 模型

Hill-climbing MAI models for GitHub Copilot and Excel

AI 导读

Microsoft 宣布将 MAI-Code-1-Flash 通过 hill-climbing 方法 specialize 到 GitHub Copilot 和 Excel 两个场景。

推荐理由

官方披露了同一 checkpoint 经 harness 内后训练扩展到 Excel 的路径和成本细节,展示了从智能体编码到知识工作迁移的做法。

正文 · 原文

Models

Superintelligence team

July 23, 2026

Models

Superintelligence team

Better models, fewer parameters, less tokens

At Build in June, we introduced our hill–climbing machine, our integrated data, model, and harness flywheel. Today we are excited to share two examples inside Microsoft: MAI models specialized for agentic workloads in GitHub Copilot and Excel.

Early results are promising. In our live product deployment, we see that our MAI model deployed in Excel is on par with GPT-5.6 for the most common tasks while being more cost-efficient.

The figure below shows how MAI-Code-1-Flash, post-trained within the GitHub Copilot harness, was used as the starting checkpoint to climb on Excel evaluations, resulting in two highly efficient, specialized models.

Line graph showing pass rates on SWE Bench (VS Code and Base) across model checkpoints, with data points labeled "Code" and "Excel." Pass rates increase over checkpoints, starting below 72% and peaking at 86%.

MAI-Code-1-Flash in GitHub Copilot

Since launching MAI-Code-1-Flash in GitHub Copilot in June, millions of developers have been using it for their day-to-day work, where it’s outperforming other similarly sized models while using fewer tokens.

  • It has an approximately 10% higher code accept rate than GPT 5.4 Mini and Claude Haiku 4.5 in VS Code.
  • Developers were 6% more likely to return across multiple days than with GPT 5.4 Mini and 11% more likely than with Claude Haiku 4.5.
  • It has 10% lower median token usage than GPT-5.4 mini and Claude Haiku 4.5, with more user-initiated turns.

MAI model live in Excel

Excel offered a test for whether the capabilities built into MAI-Code-1-Flash could transfer beyond the domain they were trained for, moving from agentic coding to agentic knowledge work. To do so, we further trained our MAI-Code-1-Flash checkpoint in an Excel reinforcement learning environment to learn about tools and knowledge workflows in spreadsheets. The result is a model with a command of Excel workflows that is more efficient and less expensive to run.

Flowchart titled "The Excel climb" showing an Excel RL Environment with action, review, execute, and update steps, linked to input, output, grader, and updated model weights.

User feedback from production traffic indicates that the quality of the MAI model in Excel is on par with GPT-5.6 for the most common tasks. In addition to the direct model cost savings, this smaller, more efficient model can be served on both Nvidia H100 and A100 class GPUs rather than requiring only the latest-generation accelerators, which significantly lowers the cost of deployment for Microsoft.

Training agentic models from inside the product stack

These results point toward a broader strategy. By having access to the entire product stack—the model, the harness that runs it, the agents, and product-specific evaluations—we can hill-climb to train efficient, powerful models capable of tasks previously handled by larger, more expensive ones.

Beyond GitHub Copilot and Excel, we’re currently extending this hill-climbing approach to train efficient models across Microsoft’s family of agentic products: Copilot Chat, Outlook, PowerPoint, and more.

Related Stories

来源:Microsoft AI:官方博客(网页) · microsoft.ai