跳到正文
原文
Rohan Paul· @rohanpaul_ai · X·· 3 小时前AI 评分66
AI 导读

Google 推出旗舰模型 Gemini 4 Argon,在多数基准上超过 GPT-6 Astra 和 Claude Opus 5.5,其中 Harvey's Legal Agent Benchmark 得分 19.6%,明显高于对比模型的 6.7% 等成绩。

正文

MASSIVE reveal from Google.

Its new flagship, Gemini 4 Argon, outscores GPT-6 Astra and Claude Opus 5.5 on most benchmarks.

- beats GPT-6 Astra and Claude Opus 5.5 on some super important industry benchmarks.

- its widest lead in legal work, 19.6% on Harvey's Legal Agent Benchmark against 6.7% for Anthropic's Claude Fable 5.1.

- output limit jumps from 64K to 1M tokens, an industry-leading ceiling,

- Only 3 groups have it today. the first is Google's own staff, vetted cyber defenders such as government agencies and security companies and trusted testers giving Google feedback.

- Inside Google, Argon agents freed over 300 TiB of data-center memory, with 500 TiB to 1 PiB of total savings estimated, and made a Rust port of the libgav1 video decoder 2.7x faster by replacing 32K lines of SIMD code.

引用Sundar Pichai@sundarpichai
Lots of discussion out there about our next model(!), so I wanted to give an early look as soon as possible. Introducing Gemini 4 Argon! It shows frontier performance in complex workflows, cyber defense and software engineering. Teams are using it extensively at Google, from coding to quantum computing, great feedback. Here’s a look at the benchmarks:
在 X 查看被引用的帖子

来源:Rohan Paul · x.com