Grok 4.7 是我们在编程和知识工作方面能力最强的模型。它在处理困难任务时能持续工作更长时间,更仔细地检查自己的工作,并配备了迄今为止我们校准最完善的防护措施。其服务价格和速度与 Grok 4.6 相同,在同级别中极具竞争力。
在 CursorBench 4.0 上——该基准侧重考验运行时间更长的编程任务——Grok 4.7 在性价比方面处于前沿水平。
模型改进
与 Grok 4.6 相比,Grok 4.7 采用了全新的、规模更大的基础模型。它在更难的任务组合上进行了更长时间的强化学习训练,训练权重偏向于那些需要数小时才能完成的问题。该模型更擅长验证自己的工作,并能管理更长的上下文。我们还训练 Grok 4.7 原生理解 Grok Bot 框架,使其在对话任务和通用知识工作方面表现更佳。
| Grok 4.7xhigh | Grok 4.6high | GPT-5.6 Solmax | Fable 5.1max | |
|---|---|---|---|---|
| 输入 token 价格,每百万美元 | $2 | $2 | $4 | $10 |
| 输出 token 价格,每百万美元 | $6 | $6 | $20 | $50 |
| 软件工程CursorBench 4.0 | 46.3% | 40.4% | 41.7% | 51.8% |
| 软件工程DeepSWE v1.1 | 71.0%* | 65.2% | 72.7% | 70.0% |
| 多小时办公任务AA Briefcase v1.1 | 1,657 | 1,546 | 1,487 | 1,678 |
| 多小时终端工作 Terminal-Bench 4.0 | 38.0% | 20.3% | 37.3% | 57.9% |
| 法律工作 Harvey Legal Agent Benchmark | 19.6% | 15.8% | 2.5% | 6.7% |
| 临床推理 HealthBench Professional | 56.7% | 48.5% | 60.5% | 62.1% |
| 电气工程 EEBench | 64.0% | 53.0% | 39.4% | 56.4% |
* 高投入
Grok 4.7 更擅长创建文档和演示文稿。在 GDPval 和 AA Briefcase 中,AI 被要求处理律师、护士和金融分析师等专业人士所做的工作。Grok 4.7 在这两项基准测试上均优于 Grok 4.6,并与其他前沿模型表现相当。
专业知识工作
GDPval
安全与网络安全
Grok 4.7 采用了全新的安全防护栈构建。它在我们测试过的模型中,在拒答和抗越狱方面表现最强。在网络安全和生物工作等两用领域,它在良性任务的实用性以及危险任务的安全拒答两方面均处于领先,并以 62.4% 的成绩位居 LatchBio 生物安全基准测试榜首。
Grok 4.7 在强大的网络防御能力与对合法用途的低拒绝率之间取得了平衡。在我们用于评估高风险和恶意网络任务的基准 HackerBench v0.3 上,它展现出最高的安全性,仅放行 3.3% 的高风险两用提示词,同时极少拦截合法的安全工作。我们还已开始向部分网络安全合作伙伴提供仅限邀请的 Grok 4.7 红队能力访问权限,用于防御研究。
定价与可用性
Grok 4.7 即日起可在 Cursor 和 Grok Build 中使用。它也可通过 Grok API、第三方编程工具链以及模型路由器和云平台使用。
该模型的定价为每百万输入 token 起价 $2,每百万输出 token 起价 $6。我们还提供一个快速版本,输出速度翻倍,价格也翻倍。
控制台
创建 API 密钥
docs.x.ai
阅读文档
在 Grok Build 中免费试用
立即前往 x.ai/build 开始使用。
Grok 4.7 is our most capable model for coding and knowledge work. It works longer on difficult tasks, checks its own work more carefully, and comes with our best-calibrated safeguards to date. Served at the same price and speed as Grok 4.6, it is highly competitive in its class.
On CursorBench 4.0, which stresses longer-running coding tasks, Grok 4.7 is at the frontier in price-performance.
Model Improvements
Grok 4.7 uses a new, larger base model compared to Grok 4.6. It was trained with a longer reinforcement learning run on a harder mix of tasks, weighted toward problems that take many hours to complete. The model is better at verifying its own work and managing longer context. We also trained Grok 4.7 to natively understand the Grok Bot harness, making it better at conversational tasks and general knowledge work.
| Grok 4.7xhigh | Grok 4.6high | GPT-5.6 Solmax | Fable 5.1max | |
|---|---|---|---|---|
| Input token price, $ per million | $2 | $2 | $4 | $10 |
| Output token price, $ per million | $6 | $6 | $20 | $50 |
| Software engineeringCursorBench 4.0 | 46.3% | 40.4% | 41.7% | 51.8% |
| Software engineeringDeepSWE v1.1 | 71.0%* | 65.2% | 72.7% | 70.0% |
| Multi-hour office workAA Briefcase v1.1 | 1,657 | 1,546 | 1,487 | 1,678 |
| Multi-hour terminal workTerminal-Bench 4.0 | 38.0% | 20.3% | 37.3% | 57.9% |
| Legal workHarvey Legal Agent Benchmark | 19.6% | 15.8% | 2.5% | 6.7% |
| Clinical reasoningHealthBench Professional | 56.7% | 48.5% | 60.5% | 62.1% |
| Electrical engineeringEEBench | 64.0% | 53.0% | 39.4% | 56.4% |
* high effort
Grok 4.7 is better at creating documents and presentations. In GDPval and AA Briefcase, AI is asked to work on tasks done by professionals such as lawyers, nurses, and financial analysts. Grok 4.7 improves upon Grok 4.6 on both benchmarks and performs comparably to other frontier models.
Professional knowledge work
GDPval
Safety & Cybersecurity
Grok 4.7 was built with an entirely new safeguard stack. It is the strongest model we’ve tested on refusals and jailbreak resistance. In dual-use domains like cybersecurity and biological work, it leads on both utility for benign tasks and safe refusal on dangerous ones, topping LatchBio’s biosafety benchmark at 62.4%.
Grok 4.7 balances strong cyber defense capabilities with low refusal rates for legitimate use. It shows the highest safety on HackerBench v0.3, our benchmark for risky and malicious cyber tasks, allowing only 3.3% of risky dual-use prompts through while rarely blocking legitimate security work. We’ve also started giving select cybersecurity partners invite-only access to Grok 4.7’s red-team capabilities for defense research.
Pricing and availability
Grok 4.7 is available today in Cursor and Grok Build. It is also available through the Grok API, third-party coding harnesses, and model routers and cloud platforms.
The model is priced starting at $2 per million input tokens and $6 per million output tokens. We also serve a fast variant with twice the output speed at twice the price.
Console
Create an API key
docs.x.ai
Read the docs
Try it in Grok Build for free
Get started today at x.ai/build.