我们提出 MaxProof,这是一种面向 MiniMax-M3 系列竞赛级数学证明的群体级测试时扩展框架。M3 首先训练了三种面向证明的能力——证明生成、证明验证以及基于批判的证明修复——并采用一种专为低误报率设计的纵深防御生成式验证器。
这些能力被合并到单个发布的 M3 模型中。在测试时,MaxProof 将该模型同时用作生成器、验证器、精炼器和排序器,在候选证明群体中进行搜索,并通过锦标赛选择返回最终证明。借助 MaxProof 的测试时扩展,M3 模型在 IMO 2025 上达到 35/42,在 USAMO 2026 上达到 36/42,两项成绩均超过人类金牌线。
We present MaxProof, a population-level test-time scaling framework for competition-level mathematical proof in the MiniMax-M3 series. M3 first trains three proof-oriented capabilities -- proof generation, proof verification, and critique-conditioned proof repair -- using a defense-in-depth generative verifier engineered for low false-positive rate. These capabilities are merged into a single released M3 model. At test time, MaxProof treats the model as a generator, verifier, refiner, and ranker, searches over a population of candidate proofs, and returns one final proof through tournament selection. With MaxProof test-time scaling, the M3 model reaches 35/42 on IMO 2025 and 36/42 on USAMO 2026, exceeding the human gold-medal threshold on both.