跳到正文
Artificial Analysis 完整文章·· 1 天前精选AI 评分66

Mistral 发布 Mistral Large 4,Artificial Analysis 评测称其为美中之外最智能模型

Mistral has released Mistral Large 4, making France home to the most intelligent model outside the US and China

AI 导读

Mistral 发布 Mistral Large 4(Research Public Preview),在 Artificial Analysis Intelligence Index 得分 38,与 GPT-6 Luna(max, 38)相当,为美中之外最智能模型,计划 10 月底开源 1T 参数(49B 激活)权重。

推荐理由

原文给出 Intelligence Index、Cyber Index、定价与上下文窗口等具体数据,读者可据此评估 Mistral Large 4 的真实水平与成本。

正文 · 原文

See model page

Mistral has released Mistral Large 4, scoring 38 on the Artificial Analysis Intelligence Index; France is back to having the most intelligent model from outside the US and China

Mistral has released Mistral Large 4 in Research Public Preview, with plans to release the weights of the 1T parameter (49B active) model at the end of October. It achieves 38 on the Artificial Analysis Intelligence Index, comparable to GPT-6 Luna (max, 38) and DeepSeek V4.1 Flash (max, 39). It also achieves 50 on the Artificial Analysis Cyber Index, level with GLM-5.3-Flash and ahead of models such as Kimi K3 and DeepSeek V4.1 Flash (max).

Key benchmarking results for Mistral Large 4 Preview:

➤ Most intelligent model from outside the US and China: Mistral Large 4 Preview scores 38 on the Intelligence Index, comparable to DeepSeek V4.1 Flash (max, 39) and GPT-6 Luna (max, 38). This makes it the most intelligent model from outside the US and China, ahead of countries such as South Korea and the United Arab Emirates

➤ Level with GLM-5.3-Flash on cyber defense capability: Mistral Large 4 Preview scores 50 on the Artificial Analysis Cyber Index, level with GLM-5.3-Flash (50) and behind MiMo-V2.6-Pro (56). Once its weights are released, it will rank among the top three open weights models on the Cyber Index. Its strongest result is on CyberGym-E2E-AA, where it scores 82%, ahead of MiMo-V2.6-Pro (79%) and GPT-6 Luna (max, 78%)

➤ Over 4x the Cost per Task of similar-intelligence open weights models: Mistral Large 4 Preview costs $1.13 per Intelligence Index task with standard pricing of $1.36/$4.18 per 1M input/output tokens, with $0.14 per 1M cached input tokens. For the first two weeks, Mistral Large 4 Preview will be served at a 50% launch discount, bringing its Cost per Task down to $0.57. This is still more costly than GLM-5.3-Flash ($0.25) and DeepSeek V4.1 Flash (max, $0.27)

➤ Strong document and image reasoning: Mistral Large 4 Preview scores 19% on GDP.pdf, on par with MiMo-V2.6-Pro (19%) and behind Kimi K3 (22%). This is an 18-point improvement from Mistral Large 3, partly driven by improvements in their API, which now accepts 100 images per request, up from 8 for previous Mistral models

Key model details:

➤ Context Window: 512k tokens

➤ Multimodality: Text and image input, with text output

➤ Pricing: $1.36/$4.18 per 1M input/output tokens ($0.14 per 1M cached input tokens), with 50% off for the first two weeks ($0.68/$2.09)

➤ Availability: Research Public Preview on Mistral's API, with open weights planned for the end of October

来源:Artificial Analysis 完整文章 · artificialanalysis.ai