Grok 4.7 by @SpaceXAI and @elonmusk is now in the Agent Arena!
Your votes drive the @arena leaderboards, head over and bring your toughest prompts.
In Agent Arena, we measure models on millions of real-world, long-horizon agentic tasks from a global community of users. Models can access web search, filesystem, and terminal tools to complete complex workflows. The leaderboard measures model performance on outcomes relative to the average model using a causal tracing methodology.
In addition to Agent Arena, Grok 4.7 is in Battle Mode for: Text, Vision, Code, and Document.