Arena.ai· @arena · X·· 2 小时前精选AI 评分71
AI 导读
Arena 宣布 Mistral Large 4 进入 Agent Arena 前十五实验室,基于超 5000 个真实智能体会话,该预览版净改进分为 -6.6%,总排名第 43,比前代 Mistral Medium 3.5(-12.60%)高出 11 位。
推荐理由
Arena 用 Agent Arena 实测数据给出 Mistral Large 4 的排名与分数,可对照其官方发布信息核实模型实际表现。
正文 · 原文
Mistral Large 4 by @MistralAI has landed in the Agent Arena top 15 labs!
Across +5K real-world agentic sessions, this preview model records a -6.6% net improvement score. It is 11 rankings above the previous variant, Mistral Medium 3.5 (-12.60%).
Mistral Large 4 is #43 overall in the Agent Arena. Open weights are expected at the end of October, but at its current score Mistral Large 4 would rank #13 among open models.
@MistralAI is also the only European lab in the Agent Arena top 15. Congrats to the team!
Meet Mistral Large 4, aka Le Chonk. • 1T parameters, natively multimodal. 49B active. It is the best open weights model from US or Europe on aggregated benchmarks. • State-of-the-art on critical workloads, including cyber defense, manufacturing and finance and it surpasses closed frontier models on visual grounding. • Forged in Europe end-to-end and is deployable from Europe via our own Mistral Cloud infrastructure. • Available to all via API today. Working with cybersecurity partners privately. Open weights release end of October.在 X 查看被引用的帖子
来源:Arena.ai · x.com