🚨 AI News | TestingCatalog· @testingcatalog · X·· 1 小时前AI 评分57
AI 导读
Google 发布 Gemini 4 Argon 前沿模型,面向真实软件工程、法律与金融等企业知识工作和网络防御的复杂工作流,DeepSWE v1.1 得分 77.9% 为新 SOTA。作者称其在多项基准上优于 GPT-6 Astra、Claude Opus 5.5 和 Claude Fable 5.1,即将首先向付费 API 客户和 Google AI Ultra 订阅用户推出。Sundar Pichai 称该模型在复杂工作流、网络防御和软件工程方面表现出前沿性能,团队已在 Google 内部从编码到量子计算广泛使用。
正文
BREAKING 🔥: Google announced Gemini 4 Argon, a new frontier model for "complex workflows across real-world software engineering, enterprise knowledge work like legal and finance, and cyber defense."
77.9% score on DeepSWE v1.1 is a new SOTA. It performs better then GPT-6 Astra, Opus 5.5 and Fable 5.1 across many benchmarks.
Rolling out soon starting with paid API customers and Google AI Ultra subscribers.
Soon! 👀
Lots of discussion out there about our next model(!), so I wanted to give an early look as soon as possible. Introducing Gemini 4 Argon! It shows frontier performance in complex workflows, cyber defense and software engineering. Teams are using it extensively at Google, from coding to quantum computing, great feedback. Here’s a look at the benchmarks:在 X 查看被引用的帖子
来源:🚨 AI News | TestingCatalog · x.com