All News
xaigrokmodel-releasebenchmarksartificial-analysisai-pricing

xAI ships Grok 4.7 at half the frontier's price, and 12 points of disagreement on its coding

Grok 4.7 undercuts the frontier on price and speed, but the same coding benchmark reads 38% at xAI and 26% at the independent evaluator Artificial Analysis.

Vlad MakarovVlad Makarovreviewed and published
5 min read
xAI ships Grok 4.7 at half the frontier's price, and 12 points of disagreement on its coding

xAI released Grok 4.7 on September 21, 2026, pitching it in its own launch post as "our most powerful model for coding and knowledge work. Twice as fast, at half the price of comparable models." The price half checks out: $2 per million input tokens and $6 per million output, unchanged from Grok 4.6, against $4 and $20 for GPT-5.6 Sol and $10 and $50 for Claude Fable 5.1. The capability half is murkier: the same coding benchmark reads 38.0% in xAI's table and 26% at the independent evaluator Artificial Analysis. Twelve points, one test, two provenances.

What xAI says it changed

The company describes a new, larger base model than Grok 4.6, trained with "a longer reinforcement learning run on a harder mix of tasks, weighted toward problems that take many hours to complete." It says the model verifies its own work more carefully and handles longer context, and was trained to "natively understand the Grok Bot harness." Grok 4.7 ships in Cursor, Grok Build, the Grok API, third-party coding harnesses and cloud platforms. No parameter count, corpus or RL recipe was published; Artificial Analysis notes SpaceXAI — the brand xAI now uses on its site — has not disclosed model size.

The vendor's numbers, and who measured them

Every figure below is xAI's own account, run on xAI's chosen harnesses. None of it has been independently audited.

MetricGrok 4.7 xHighGrok 4.6 HighGPT-5.6 Sol MaxFable 5.1 Max
Input price per M$2$2$4$10
Output price per M$6$6$20$50
CursorBench 4.046.3%40.4%41.7%51.8%
DeepSWE v1.171.0%*65.2%72.7%70.0%
EEBench64.0%53.0%39.4%56.4%
AA Briefcase v1.11,6571,5461,4871,678
Terminal-Bench 4.038.0%20.3%37.3%57.9%
Harvey Legal Agent19.6%15.8%2.5%6.7%
HealthBench Professional56.7%48.5%60.5%62.1%

The asterisk marks a high-effort Grok 4.7 result. Rival scores inside a vendor table are vendor-measured too: the Fable 5.1 and GPT-5.6 Sol columns come from xAI, not Anthropic or OpenAI.

Even as xAI tells it, the picture is mixed. Grok 4.7 wins EEBench, but on xAI's own page it loses Terminal-Bench 4.0 to Fable 5.1 by 19.9 points and CursorBench to Fable 5.1's 51.8. The same post's GDPval chart puts Fable 5.1 (max) at 1,735 Elo, ahead of Grok 4.7 (xhigh) at 1,695 and GPT-6 Astra (max) at 1,542.

One benchmark, two numbers twelve points apart

Terminal-Bench 4.0 is a 66-task agentic terminal benchmark built by the Laude Institute, Stanford researchers and open-source contributors. The 4.0 release recalibrated compute and time allowances and removed eight tasks that were saturated, refusal-prone or publicly solved. Artificial Analysis runs all 66 tasks with the mini-swe-agent harness and reports pass@1 averaged over three repeats. xAI's 38.0% row is labelled "xHigh," an effort setting on its own stack, and the launch post does not name the harness behind it.

That method difference is the likeliest home for the twelve points, and the reason not to declare a winner: a different scaffold, effort budget and repeat policy routinely move agentic scores by double digits. Neither number is dishonest, and neither is published in a form a reader can reconcile — the same shape of dispute we covered in the ARC-AGI-3 harness gap.

The independent reading is harsher elsewhere. On the Artificial Analysis Intelligence Index (v4.3.2, ten benchmarks) Grok 4.7 scores 46 and lands mid-pack, with Fable 5.1 and GPT-6 leading at 53 each; AA notes its two highest reasoning levels "appear to perform about the same," which undercuts the top of the effort range. On Terminal-Bench, the cheaper DeepSeek V4.1 Flash scores 27%, a point above Grok 4.7.

'Twice as fast at half the price' is a price claim

The headline is a ratio against comparables, delivered as price-performance rather than a speed record. The $2/$6 pricing is the news; a fast variant is served at twice the output speed for twice the price. Artificial Analysis measures the default xhigh configuration at 39.3 output tokens per second — the low end for reasoning models in its tier, where the median is 72.5 — and calls the model "notably slow and very verbose," generating 240M output tokens on the Intelligence Index against a median of 94M. Time to first token, 0.88 seconds, is genuinely competitive. The cheap configuration the benchmarks were run on is not the fast one, so "twice as fast" and the published scores describe different products.

The safety claim is a vendor benchmark

xAI says the model carries an "entirely new safeguard stack," is the strongest it has tested on refusals and jailbreak resistance, tops LatchBio's biosafety benchmark at 62.4%, and allows only 3.3% of risky dual-use prompts through on HackerBench v0.3 — which xAI itself wrote. It has also begun giving select cybersecurity partners invite-only access to Grok 4.7's red-team capabilities for defense research. A vendor reporting strong results on its own test is a claim about the test, not an audit.

What the thread actually argues about

Hacker News item 49788838 held roughly 383 points and about 327 comments some four hours after launch, and almost none of the discussion is about Terminal-Bench. The top comment claims Grok 4.7 "has 40% more weights than Grok 4.6" — community chatter, not specification, since xAI has published no parameter count. The same comment says the release slipped "almost two weeks past the original date" and speculates about a rumored next Anthropic release; both are speculation. Another cluster prefers Grok's plainer English to what one commenter called "Claudish" output, with a side debate about prompting models toward ASD-STE100 Simple Technical English. Reddit trend data was thin, so the Hacker News thread carries the reception.

What would settle it

Three things, none exotic. xAI could name the harness, effort setting and repeat policy behind each row, or rerun them on a public harness so the readings become comparable. Both parties could publish per-task pass rates on Terminal-Bench 4.0, which would locate the disagreement in specific tasks rather than an aggregate. And someone independent could re-run the fast variant at matched effort, since speed and capability are sold together but measured apart. Until then, treat "twice as fast at half the price" as a statement about what you pay. Grok 4.7 is a credible buy at its price; it is not a capability lead, and nobody has shown that it is.

Related Articles

Scroll down

to load the next article