小米发布 MiMo-V2.6 系列,Pro 登顶开源模型榜,Anthropic 指其借 Claude 蒸馏数据

The Decoder:AI News(RSS)·2026-09-22 19:42·13分钟前·Maximilian Schreiner
AI 导读

小米发布 MiMo-V2.6 系列,旗舰 MiMo-V2.6-Pro 在 Artificial Analysis Intelligence Index 得 46 分,居开源模型之首,价格为每百万输入 token $0.435、输出 $0.87。

The Decoder:AI News(RSS)
精选
76AI 编辑部评分,满分 100

小米发布 MiMo-V2.6 系列,Pro 登顶开源模型榜,Anthropic 指其借 Claude 蒸馏数据

2026-09-22 19:42· 13分钟前· Maximilian Schreiner
AI 导读

小米发布 MiMo-V2.6 系列,旗舰 MiMo-V2.6-Pro 在 Artificial Analysis Intelligence Index 得 46 分,居开源模型之首,价格为每百万输入 token $0.435、输出 $0.87。

推荐理由

原文同时给出 MiMo-V2.6-Pro 的成绩、价格、训练成本和 Anthropic 指控细节,读者可以对照评估其开放策略与争议。

Image description

Xiaomi

Key Points

  • Xiaomi's new MiMo-V2.6-Pro model leads current rankings of open AI models while costing far less than the competition.
  • Xiaomi achieved the performance jump through expanded reinforcement learning, and the company released its tools and training tasks openly.
  • At the same time, Anthropic accuses Xiaomi of improperly siphoning off data through the Claude model to train its own model lineup.

Xiaomi has released its MiMo-V2.6 lineup, and the flagship model tops the charts among openly available models while costing a fraction of the competition per task. The gains come from a massively expanded round of reinforcement learning, though the company also stands accused of borrowing from Anthropic's Claude.

According to Xiaomi, the larger of the two new models, MiMo-V2.6-Pro, scores 46 points on the Intelligence Index from analysis firm Artificial Analysis. That makes it the strongest openly available AI model right now, ahead of rivals like Kimi K3 and Qwen. The real kicker is the price: $0.435 per million input tokens and $0.87 per million output tokens.

By Artificial Analysis's math, a single test task costs only about $0.13, a fraction of what similarly capable models charge. That puts the model on what's called the Pareto frontier of intelligence and cost.

Image: Artificial Analysis
Image: Artificial Analysis

Pro is a mixture-of-experts model with 1.02 trillion parameters, only 42 billion of which are active per request. Alongside it sits the smaller, more efficient MiMo-V2.6-Flash.

The gains come from reinforcement learning

Xiaomi credits the jump in performance to heavily expanded reinforcement learning (RL), meaning training through trial, feedback, and reward. The company scaled this phase along three axes: more data per training step, more varied task environments, and more compute for grading the solutions.

The run took less than six days, Xiaomi says, and cost about $2.62 million for Pro and $0.85 million for Flash. On the DeepSWE coding test, Pro's score climbed from 58.4 to 72.6, while Flash rose from 48.8 to 65.7.

To keep training stable at this scale, Xiaomi froze the model's internal distribution mechanism and added several layers of protection against "reward hacking," the tricks a model uses to game rewards without actually solving the task.

About 7,000 tasks with automatic graders

Along with the models, Xiaomi is shipping an especially fast variant called Pro-UltraSpeed, with up to 20 times the output speed. What stands out most is that the company is opening up its RL toolkit, including the technical report, the full training framework, a smaller model for further training, and about 7,000 ready-made training tasks with automatic graders for software development, cybersecurity, office work, and web design, plus roughly 1,000 tasks for music composition.

The tasks come from a mix of sources. Some of the code comes from real GitHub pull requests by employees and user queries, while other task descriptions are generated by a language model. The cyber tasks draw on OSS-Fuzz, a collection of tens of thousands of real software vulnerabilities, and the office environments are rebuilt synthetically.

Pointedly open, and in Anthropic's crosshairs

This show of openness sits in sharp contrast to accusations Anthropic raised just two weeks earlier. In its threat intelligence report, Anthropic examined cases of Claude abuse discovered between December 2025 and August 2026, and named seven Chinese labs tied to campaigns against the model: Alibaba, Moonshot AI, DeepSeek, Zhipu, Xiaomi, MiniMax, and SenseTime.

All told, the labs are said to have generated about 190 million exchanges to siphon off Claude's capabilities for training their own models, a technique Anthropic calls illegal distillation.

Xiaomi shows up by name. In a case tagged GTG-16008, Anthropic tracked more than 400,000 exchanges over 20 days in March and April 2026, in which Xiaomi passed user conversations and coding sessions from its own MiMo models through OpenClaw and OpenCode to Claude, aiming to enrich training data for future models. The report offers almost no documentation of where the earlier training and teacher data for the internal distillation of teacher models came from.

Put another way, the report claims Xiaomi recorded user conversations with its MiMo models and then fed them into Claude to extract training data, the same data that has now pushed it to the top of the open-model rankings.

MiMo

来源:The Decoder:AI News(RSS)· the-decoder.com