SpaceXAI 发布了 Grok 4.7,这是其面向编程、智能体任务和知识工作的新旗舰模型。Grok 4.7 基于更大的基础模型和更长的强化学习训练构建。它的价格和速度仍与 Grok 4.6 保持一致。
它可以部署吗? 可以,作为托管模型部署。你今天就可以通过 xAI API、Cursor、Grok Build、OpenRouter、Vercel 和 Cloudflare 调用 grok-4.7。
底层有哪些变化
SpaceXAI 列出了相较 Grok 4.6 的 4 项变化:
- 一个全新的、更大的基础模型:Grok 4.7 没有复用 Grok 4.6 的基础模型。
- 在更难任务上进行更长的强化学习训练:任务组合更偏向那些需要数小时才能完成的问题。
- 更好的自我验证和长上下文处理能力:该公司表示,该模型会更仔细地检查自己的工作。
- 原生支持 Grok Bot harness:它经过训练,能够理解用于对话和知识工作的 Grok Bot harness。
这份开发者文档列出了 API 规格:
| 属性 | 值 |
|---|---|
| 模型名称 | grok-4.7 |
| 上下文窗口 | 500,000 tokens |
| 知识截止日期 | 2026 年 5 月 |
| 模态 | 文本和图像输入,文本输出 |
| 推理强度 | low、medium、high(默认)、xhigh |
| API | Responses API、Chat Completions |
| 工具 | 函数调用、网页搜索、X 搜索、代码执行 |
基准测试
发布对比表将 xHigh 努力档的 Grok 4.7 与 Grok 4.6 High、GPT-5.6 Sol Max 和 Fable 5.1 Max 进行了比较。Grok 4.7 的 DeepSWE 分数是在 high 努力档下运行的。所有分数均为厂商自报。
| 基准测试 | Grok 4.7 xHigh | Grok 4.6 High | GPT-5.6 Sol Max | Fable 5.1 Max |
|---|---|---|---|---|
| 输入价格($/M) | $2 | $2 | $4 | $10 |
| 输出价格($/M) | $6 | $6 | $20 | $50 |
| CursorBench 4.0 | 46.3% | 40.4% | 41.7% | 51.8% |
| DeepSWE v1.1 | 71.0%* | 65.2% | 72.7% | 70.0% |
| EEBench | 64.0% | 53.0% | 39.4% | 56.4% |
| AA Briefcase v1.1 | 1,657 | 1,546 | 1,487 | 1,678 |
| Terminal-Bench 4.0 | 38.0% | 20.3% | 37.3% | 57.9% |
| Harvey 法律智能体基准 | 19.6% | 15.8% | 2.5% | 6.7% |
| HealthBench Professional | 56.7% | 48.5% | 60.5% | 62.1% |
*高投入
Grok 4.7 在每一行上都优于 Grok 4.6。提升最大的是 Terminal-Bench 4.0,从 20.3% 升至 38.0%。EEBench 上升 11 个百分点至 64.0%,为表中最高分。在 Harvey 的法律智能体基准上,Grok 4.7 得分为 19.6%,而 Fable 5.1 Max 为 6.7%。
Grok 4.7 并非全面领先。Fable 5.1 Max 在 7 项基准测试中拿下 4 项第一,其中包括 57.9% 的 Terminal-Bench 得分。GPT-5.6 Sol Max 则保持着 DeepSWE v1.1 的最高成绩,达到 72.7%。
价格是另一个维度。Fable 5.1 Max 的输入成本高出 5 倍,输出成本高出约 8.3 倍。GPT-5.6 Sol Max 的输入成本高出 2 倍,输出成本高出约 3.3 倍。在 CursorBench 4.0 的每任务成本图表上,SpaceXAI 将 Grok 4.7 置于性价比的前沿位置。
在测试专业知识工作的 GDPval 上,Grok 4.7 xhigh 得分为 1,695 Elo。这较 Grok 4.6 high 的 1,605 有所提升。Fable 5.1 max 以 1,735 领先,GPT-6 Astra max 得分为 1,542。SpaceXAI 还表示,Grok 4.7 在创建文档和演示文稿方面表现更佳。
安全与网络安全
Grok 4.7 搭载了一套全新的防护体系。SpaceXAI 称其是所测试过的在拒答和抗越狱方面最强的模型。它以 62.4% 的成绩登顶 LatchBio 的生物安全基准测试。
在 HackerBench v0.3 上——这是 SpaceXAI 自研的针对高风险和恶意网络任务的基准测试——该模型放行了 3.3% 的高风险两用提示词。该公司表示,它很少拦截合法的安全工作。部分网络安全合作伙伴现在可通过邀请制获取其红队能力,用于防御研究。
定价与可用性
Grok 4.7 的价格为每百万输入 token 2 美元,每百万输出 token 6 美元。它已在所有套餐的 Cursor 中可用,并且是 Grok Build 中的默认模型。它还通过 Grok API、OpenRouter、Vercel 和 Cloudflare 提供服务。
Grok 4.7 Fast 是同一模型运行在更快的基础设施上,输出速度翻倍,价格也翻倍。文档称它仅在 Cursor 和 Grok Build 中运行,不在公开的 xAI API 上提供。它也被排除在 Grok Build 的免费套餐之外。
A美国区域端点在https://us.api.x.ai/v1将推理保留在美国境内,溢价 10%。SpaceXAI 建议设置一个prompt_cache_key以实现可靠的缓存命中。
import os
from xai_sdk import Client
from xai_sdk.chat import user
client = Client(api_key=os.getenv("XAI_API_KEY"))
chat = client.chat.create(model="grok-4.7")
chat.append(user("Explain this repo."))
print(chat.sample().content) 交互式讲解
核心要点
- Grok 4.7 维持 Grok 4.6 的定价,即每百万 tokens 输入 $2、输出 $6。
- 它在 EEBench(64.0%)和 Harvey Legal Agent Benchmark(19.6%)的发布榜单上领先。
- Fable 5.1 Max 仍在 7 项基准中的 4 项上领先,包括 Terminal-Bench 4.0。
- 该 API 提供 500,000-token 的上下文窗口,以及最高至 xhigh 的 4 个推理等级。
- 在 SpaceXAI 的 HackerBench v0.3 上,只有 3.3% 的高风险两用提示词通过了测试。
查看 公告、文档以及 X 帖子。所有功劳都归功于该项目的研究人员。另外,欢迎在 Twitter 上关注我们,别忘了加入我们拥有 15 万+ 成员的机器学习 SubReddit,并订阅我们的新闻通讯。等等!你在用 Telegram 吗?现在你也可以在 Telegram 上加入我们了。
SpaceXAI has released Grok 4.7, its new flagship model for coding, agentic tasks, and knowledge work. Grok 4.7 is built on a larger base model and a longer reinforcement learning run. It still ships at the same price and speed as Grok 4.6.
Is it deployable? Yes, as a hosted model. You can call grok-4.7 today through the xAI API, Cursor, Grok Build, OpenRouter, Vercel, and Cloudflare.
What Changed Under the Hood
SpaceXAI lists 4 changes over Grok 4.6:
- A new, larger base model: Grok 4.7 does not reuse the Grok 4.6 base.
- A longer RL run on harder tasks: The task mix is weighted toward problems that take many hours to complete.
- Better self-verification and long-context handling: The company says the model checks its own work more carefully.
- Native Grok Bot harness support: It was trained to understand the Grok Bot harness for conversational and knowledge work.
The developer docs list the API specs:
| Property | Value |
|---|---|
| Model name | grok-4.7 |
| Context window | 500,000 tokens |
| Knowledge cutoff | May 2026 |
| Modalities | Text and image input, text output |
| Reasoning effort | low, medium, high (default), xhigh |
| APIs | Responses API, Chat Completions |
| Tools | Function calling, web search, X search, code execution |
Benchmarks
The launch table compares Grok 4.7 at xHigh effort with Grok 4.6 High, GPT-5.6 Sol Max, and Fable 5.1 Max. The Grok 4.7 DeepSWE score was run at high effort. All scores are vendor-reported.
| Benchmark | Grok 4.7 xHigh | Grok 4.6 High | GPT-5.6 Sol Max | Fable 5.1 Max |
|---|---|---|---|---|
| Input price ($/M) | $2 | $2 | $4 | $10 |
| Output price ($/M) | $6 | $6 | $20 | $50 |
| CursorBench 4.0 | 46.3% | 40.4% | 41.7% | 51.8% |
| DeepSWE v1.1 | 71.0%* | 65.2% | 72.7% | 70.0% |
| EEBench | 64.0% | 53.0% | 39.4% | 56.4% |
| AA Briefcase v1.1 | 1,657 | 1,546 | 1,487 | 1,678 |
| Terminal-Bench 4.0 | 38.0% | 20.3% | 37.3% | 57.9% |
| Harvey Legal Agent Benchmark | 19.6% | 15.8% | 2.5% | 6.7% |
| HealthBench Professional | 56.7% | 48.5% | 60.5% | 62.1% |
*High effort
Grok 4.7 improves on Grok 4.6 in every row. The largest jump is on Terminal-Bench 4.0, from 20.3% to 38.0%. EEBench rose 11 points to 64.0%, the top score in the table. On Harvey’s legal agent benchmark, Grok 4.7 scored 19.6% against 6.7% for Fable 5.1 Max.
Grok 4.7 does not lead across the board. Fable 5.1 Max tops 4 of 7 benchmarks, including a 57.9% Terminal-Bench score. GPT-5.6 Sol Max holds the top DeepSWE v1.1 result at 72.7%.
Price is the other axis. Fable 5.1 Max costs 5x more on input and about 8.3x more on output. GPT-5.6 Sol Max costs 2x more on input and about 3.3x more on output. On a CursorBench 4.0 cost-per-task chart, SpaceXAI places Grok 4.7 at the frontier in price-performance.
On GDPval, which tests professional knowledge work, Grok 4.7 xhigh scored 1,695 Elo. That is up from 1,605 for Grok 4.6 high. Fable 5.1 max leads at 1,735, and GPT-6 Astra max scored 1,542. SpaceXAI also says Grok 4.7 is better at creating documents and presentations.
Safety and Cybersecurity
Grok 4.7 ships with an entirely new safeguard stack. SpaceXAI calls it the strongest model it has tested on refusals and jailbreak resistance. It topped LatchBio’s biosafety benchmark at 62.4%.
On HackerBench v0.3, SpaceXAI’s own benchmark for risky and malicious cyber tasks, the model let 3.3% of risky dual-use prompts through. The company says it rarely blocks legitimate security work. Select cybersecurity partners now get invite-only access to its red-team capabilities for defense research.
Pricing and Availability
Grok 4.7 costs $2 per million input tokens and $6 per million output tokens. It is available in Cursor on all plans and is the default model in Grok Build. It is also served through the Grok API, OpenRouter, Vercel, and Cloudflare.
Grok 4.7 Fast is the same model on faster infrastructure, with twice the output speed at twice the price. The docs say it runs only in Cursor and Grok Build, not on the public xAI API. It is also excluded from Grok Build’s free tier.
A US regional endpoint at https://us.api.x.ai/v1 keeps inference in the United States at a 10% premium. SpaceXAI recommends setting a prompt_cache_key for reliable cache hits.
import os
from xai_sdk import Client
from xai_sdk.chat import user
client = Client(api_key=os.getenv("XAI_API_KEY"))
chat = client.chat.create(model="grok-4.7")
chat.append(user("Explain this repo."))
print(chat.sample().content) Interactive Explainer
Key Takeaways
- Grok 4.7 keeps Grok 4.6 pricing at $2 input and $6 output per million tokens.
- It leads the launch table on EEBench (64.0%) and Harvey Legal Agent Benchmark (19.6%).
- Fable 5.1 Max still leads on 4 of 7 benchmarks, including Terminal-Bench 4.0.
- The API offers a 500,000-token context window and 4 reasoning levels up to xhigh.
- Only 3.3% of risky dual-use prompts passed on SpaceXAI’s HackerBench v0.3.
Check out the announcement, docs, and X post. All credit goes to the researcher of this project. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well.