← 返回本期xAI 发布 Grok 4.7,称其是面向编程与知识工作能力最强的模型,定价与速度与 Grok 4.6 保持一致。提升来自更大的基座模型,以及在更难、常需数小时完成的任务上做更长的强化学习训练,使模型更擅长验证自身工作与管理长上下文;它还被训练为原生理解 Grok Bot 运行框架。官方称在侧重长时编程任务的 CursorBench 4.0 上,它相对 Fable 5.1、Opus 5、GPT-5.6 Sol 与 Sonnet 5 处于性价比前沿。基准结果有输有赢:EEBench 64.0%、Harvey Legal Agent Benchmark 19.6%、AA Briefcase v1.1 为 1657,领先对照模型;但 CursorBench 4.0 为 46.3%、落后 Fable 5.1 的 51.8%,Terminal-Bench 4.0 为 37.6%、大幅落后 Fable 5.1 的 57.9%,HealthBench Professional 为 56.7%、落后 GPT-5.6 Sol 的 60.5%。GDPval 上 Grok 4.7 xhigh 的 Elo 为 1695,高于 Grok 4.6 的 1605,低于 Fable 5.1 max 的 1735。安全是核心卖点:xAI 称其为迄今测试过在拒答与抗越狱方面最强的模型,在 LatchBio 生物安全基准上以 62.4% 位居榜首,仅放行 3.3% 的高风险两用网络提示词,同时极少拦截合法的安全工作。定价为每百万输入 Token 2 美元、输出 6 美元,另有速度与价格均翻倍的快速版本;模型已在 Cursor、Grok Build、Grok API、第三方编程运行框架、模型路由器与云平台上线。

文章 · Cursor Blog

Grok 4.7 发布

原题:Introducing Grok 4.7

xAI约 3 分钟
内容摘要xAI 发布 Grok 4.7,定位编程与知识工作,价格与速度与 Grok 4.6 持平(每百万输入 2 美元、输出 6 美元),提升来自更大基座与更长强化学习,基准成绩有领先也有落后,安全防护是核心卖点。

Grok 4.7 is our most capable model for coding and knowledge work. It works longer on difficult tasks, checks its own work more carefully, and comes with our best-calibrated safeguards to date. Served at the same price and speed as Grok 4.6, it is highly competitive in its class.

55%CursorBench 4.0 score 50% 45% 40% 35% 30% 25% 20% $18 $15 $12 $9 $6 $3 $0 Average cost per task Fable 5.1Opus 5GPT-5.6 SolSonnet 5Grok 4.7

On CursorBench 4.0, which stresses longer-running coding tasks, Grok 4.7 is at the frontier in price-performance.

Model Improvements

Grok 4.7 uses a new, larger base model compared to Grok 4.6. It was trained with a longer reinforcement learning run on a harder mix of tasks, weighted toward problems that take many hours to complete. The model is better at verifying its own work and managing longer context. We also trained Grok 4.7 to natively understand the Grok Bot harness, making it better at conversational tasks and general knowledge work.

Grok 4.7 xHigh
Grok 4.6 High
GPT-5.6 Sol Max
Fable 5.1 Max
Input token price$ per million
$2
$2
$4
$10
Output token price$ per million
$6
$6
$20
$50
Software engineeringCursorBench 4.0
46.3%
40.4%
41.7%
51.8%
Software engineeringDeepSWE v1.1
71.0%*
65.2%
72.7%
70.0%
Electrical engineeringEEBench
64.0%
53.0%
39.4%
56.4%
Multi-hour office workAA Briefcase v1.1
1,657
1,546
1,487
1,678
Multi-hour terminal workTerminal-Bench 4.0
37.6%
20.3%
37.3%
57.9%
Legal workHarvey Legal Agent Benchmark
19.6%
15.8%
2.5%
6.7%
Clinical reasoningHealthBench Professional
56.7%
48.5%
60.5%
62.1%

* high effort

Token prices and benchmark scores for Grok 4.7, Grok 4.6, GPT-5.6 Sol, and Fable 5.1. Benchmarks are CursorBench 4.0, DeepSWE v1.1, EEBench, AA Briefcase v1.1, Terminal-Bench 4.0, Harvey Legal Agent Benchmark, and HealthBench Professional. An asterisk on Grok 4.7 DeepSWE marks a high-effort score.

Grok 4.7 is better at creating documents and presentations. In GDPval and AA Briefcase, AI is asked to work on tasks done by professionals such as lawyers, nurses, and financial analysts. Grok 4.7 improves upon Grok 4.6 on both benchmarks and performs comparably to other frontier models.

Professional knowledge work

GDPval

0 500 1000 1500 Elo score 1735Fable 5.1  (max) 1695 Grok 4.7 (xhigh) 1605 Grok 4.6 (high) 1542 GPT-6 Astra (max)
GDPval, AA Briefcase, and EEBench scores comparing Grok 4.7 with Grok 4.6, Fable 5.1, and GPT-6 Astra.

Safety & Cybersecurity

Grok 4.7 was built with an entirely new safeguard stack. It is the strongest model we’ve tested on refusals and jailbreak resistance. In dual-use domains like cybersecurity and biological work, it leads on both utility for benign tasks and safe refusal on dangerous ones, topping LatchBio’s biosafety benchmark at 62.4%.

Grok 4.7 balances strong cyber defense capabilities with low refusal rates for legitimate use. It shows the highest safety on HackerBench v0.3, our benchmark for risky and malicious cyber tasks, allowing only 3.3% of risky dual-use prompts through while rarely blocking legitimate security work. We’ve also started giving select cybersecurity partners invite-only access to Grok 4.7’s red-team capabilities for defense research.

Pricing and availability

Grok 4.7 is available today in Cursor and Grok Build. It is also available through the Grok API, third-party coding harnesses, and model routers and cloud platforms.

The model is priced starting at $2 per million input tokens and $6 per million output tokens. We also serve a fast variant with twice the output speed at twice the price.