跳到正文
Artificial Intelligence News·· 1 天前AI 评分71

Mistral AI 发布 Mistral Large 4 公开预览版,1 万亿参数开源权重预计 10 月底放出

Mistral AI launches Large 4 preview ahead of open-weight release

AI 导读

Mistral AI 发布原生多模态模型 Mistral Large 4 的公开预览版,总参数 1 万亿、活跃参数 490 亿,预览 API 已在 Mistral Studio 开放,权重计划 2026 年 10 月底前随架构细节和后训练方法一同发布。

正文

Mistral AI has announced a public preview of Mistral Large 4, a natively multimodal model with 1 trillion parameters and 49 billion active parameters. The preview API is available through Mistral Studio, with weights scheduled for release by the end of October 2026.

In its announcement, Mistral said it trained the model – nicknamed ‘le Chonk’ (yes, the name which started as a meme) – from scratch on 3,800 NVIDIA Grace Blackwell GPUs in its own European datacentres. The public preview runs on the same infrastructure. Its training data spans more than 160 languages, including every official language in the EU.

Mistral AI said: “We are red-teaming the model in real-world settings with cybersecurity leaders, vetted partners, and state authorities, who will access the same model with reduced moderation and expanded cyber capabilities.”

Cybersecurity tests and Mistral AI’s Large 4 model deployment plans

According to Mistral, Large 4 ranks among the top five models on the Artificial Analysis Cyber Index, which evaluates finding and fixing security flaws in real software. It reports an 82 percent score on a test requiring models to reproduce a real vulnerability in open-source software and then patch it—the highest score of any model on that test, according to the company.

Mistral also reports that the model solves 93 percent of Cybench’s 40 security-competition exercises. In internal testing, the company found it useful for malware analysis, vulnerability prioritisation, and detection-rule writing.

The company argues that provider-level refusals can obstruct legitimate vulnerability research and incident response. It plans to support private-cloud and on-premise operation through the open-weight release, allowing organisations to apply their own policies.

Separately, Mistral reports that Large 4 resisted 93.3 percent of attacks on Lakera’s public B3 AI Security Benchmark (Note that is a benchmark result, rather than evidence of protection against almost every prompt-injection attack.)

Coding agents and visual tasks

For software engineering, Mistral cites scores of 61.7 percent on DeepSWE v1.1, 59.4 percent on SWE-Atlas-QnA, and 28.3 percent on Terminal-Bench 4. These figures come from the Artificial Analysis coding index. Its combined Coding Agent Index score is 49.8 percent.

A separate blind human evaluation conducted with Surge AI asked professional annotators to rate coding outputs on a 1–5 scale with model identities hidden. Large 4 Preview ranked second among five models, scoring 3.74, behind Claude Opus 5 at 4.22.

On AutomationBench – which covers 657 business workflows across applications including Gmail, Google Sheets, Slack, and Salesforce – Mistral reports a 59.9 percent score.

The company also demonstrated visual tasks involving technical drawings, PDFs, and satellite imagery. Its reported Dense 200 visual-grounding result was 42 percent, compared with 41 percent for GPT-6-Astra in Mistral’s testing.

Reinforcement learning across training environments

Mistral said Large 4 uses the same training, customisation, and reinforcement-learning environment offered to customers through Mistral Forge.

The model’s reinforcement-learning library combines chat, scientific problem-solving, safety alignment, factuality, and long-running tool-use tasks within a shared interface. Training environments use resources including code sandboxes, web search, and external APIs, with verification through reward models, unit tests, LLM judges, and static checks.

At a reported scale of 3,000 GPUs, a training run produces roughly 33 billion tokens per day, including around 16 billion trainable completion tokens after filtering and masking. Mistral said the reinforcement-learning run behind the preview remains in progress.

The company plans to release the weights by the end of October 2026, alongside architecture details, additional benchmarks, and its post-training methodology.

See also: CyberDSA 2026 opens with Malaysia’s call for secure AI adoption

Banner for AI & Big Data Expo by TechEx events.

Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and is co-located with other leading technology events including IoT Tech Expo and the Cyber Security & Cloud Expo. Click here for more information.

AI News is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here.

来源:Artificial Intelligence News · artificialintelligence-news.com