Grok 4.7 上线 Arena 的 Agent Arena 评测

Arena.ai · @arena · X·2026-09-22 00:37·35分钟前
AI 导读

Arena 宣布 xAI 的 Grok 4.7 已加入 Agent Arena 评测,用户可投票影响排行榜。该评测基于数百万真实的长周期智能体任务,模型可使用网页搜索、文件系统和终端工具完成复杂工作流,并用因果追踪方法衡量相对表现。Grok 4.7 还进入 Battle Mode 的文本、视觉、代码和文档类别。引用方称其相较 Grok 4.6 在同价同速下有明显提升。

Arena.ai@arena
63AI 编辑部评分,满分 100

Grok 4.7 上线 Arena 的 Agent Arena 评测

2026-09-22 00:37· 35分钟前
AI 导读

Arena 宣布 xAI 的 Grok 4.7 已加入 Agent Arena 评测,用户可投票影响排行榜。该评测基于数百万真实的长周期智能体任务,模型可使用网页搜索、文件系统和终端工具完成复杂工作流,并用因果追踪方法衡量相对表现。Grok 4.7 还进入 Battle Mode 的文本、视觉、代码和文档类别。引用方称其相较 Grok 4.6 在同价同速下有明显提升。

Grok 4.7 by @SpaceXAI and @elonmusk is now in the Agent Arena!

Your votes drive the @arena leaderboards, head over and bring your toughest prompts.

In Agent Arena, we measure models on millions of real-world, long-horizon agentic tasks from a global community of users. Models can access web search, filesystem, and terminal tools to complete complex workflows. The leaderboard measures model performance on outcomes relative to the average model using a causal tracing methodology.

In addition to Agent Arena, Grok 4.7 is in Battle Mode for: Text, Vision, Code, and Document.

SpaceXAIGrok 4.7 is here. It's a notable improvement over Grok 4.6 at the same price and speed.