xAI logo

Grok 4.7

Multimodal
xAI

Grok 4.7 is SpaceXAI's frontier model for coding, agentic tasks, and knowledge work. It uses a larger base model than Grok 4.6, with longer RL on harder multi-hour tasks, stronger self-verification, and native Grok Bot harness understanding. Reasoning efforts: low, medium, high, xhigh (default high). 500K context; same price and speed class as Grok 4.6 ($2/$0.50 cached/$6 per 1M tokens under 200k prompts; ≥200k prompts bill at $4/$1/$12). Grok 4.7 Fast uses the same model on faster infrastructure at twice the token rates and twice the output speed, and is available only in Cursor and Grok Build — not as a separate public API model id (docs expose only `grok-4.7`). Pretraining knowledge cutoff June 2026, with supplemental training data through August 2026. Multimodal: text and image in, text out.

Key Specifications

Parameters
-
Context
500.0K
Release Date
September 21, 2026
Average Score
56.5%

Timeline

Key dates in the model's history
Announcement
September 21, 2026
Last Update
September 22, 2026
Today
October 5, 2026

Technical Specifications

Parameters
-
Training Tokens
-
Knowledge Cutoff
June 1, 2026
Family
-
Capabilities
MultimodalZeroEval

Pricing & Availability

Input (per 1M tokens)
$2.00
Output (per 1M tokens)
$6.00
Max Input Tokens
500.0K
Max Output Tokens
-
Supported Features
Function CallingStructured OutputCode ExecutionWeb SearchBatch InferenceFine-tuning

Benchmark Results

Model performance metrics across various tests and benchmarks

Other Tests

Specialized benchmarks
AA-Briefcase v1.1
AA Briefcase v1.1 Elo; Grok 4.7 xhigh (1657; max_score 3000). Aligned to aa-briefcase-1.1. Sources: announcement + model card PDF. • Self-reported
55.2%
BixBench MCQ
BixBench zero-shot MCQ accuracy; Grok 4.7 high (88.4%). Distinct from agentic BixBench. Sources: model card PDF. • Self-reported
88.4%
CADGenBench
CADGenBench generation split; Grok 4.7 high (44.4%) with Grok Build harness. Not BenchCAD. Sources: model card PDF. • Self-reported
44.4%
CathedralBench
CathedralBench hard-subset accuracy; Grok 4.7 xhigh (29%). Unrestricted third-party red-team cyber eval. Sources: model card PDF. • Self-reported
29.0%
CursorBench 4.0
CursorBench 4.0; Grok 4.7 xhigh (46.3%). Model card also reports 43.9% at high. Scores not comparable to CursorBench 3.2. Sources: announcement + model card PDF. • Self-reported
46.3%
CVE-Bench
CVE-Bench Reward; Grok 4.7 xhigh primary (36.6%); model card also reports 37.7% at high — not blended. Grok Build harness, unrestricted. Sources: model card PDF. • Self-reported
36.6%
CyberGym
CyberGym Mean Reproduced %; Grok 4.7 high (80.3%), unrestricted / without standard safeguards; Grok Build harness. Sources: model card PDF. • Self-reported
80.3%
DeepSWE 1.1
DeepSWE v1.1; Grok 4.7 high (71.0%; launch comparison table marks high effort with asterisk). Sources: announcement + model card PDF. • Self-reported
71.0%
EEBench
EEBench; Grok 4.7 xhigh (66.0% from model card, Grok Build harness). News comparison table lists 64.0%; catalog stores the model-card figure, not a blend. Sources: model card PDF + announcement. • Self-reported
66.0%
FrontierSWE V2
FrontierSWE V2 mean@5 (29.0%) at xhigh; Proximal Labs / Proximus harness. Not comparable to V1 dominance scores. Sources: model card PDF. • Self-reported
29.0%
GDPval (Elo)
GDPval Elo; Grok 4.7 xhigh 1695 (max_score 3000). Aligned to gdpval-elo. Sources: announcement + model card PDF. • Self-reported
56.5%
Harvey LAB (Vals)
Harvey LAB (Vals); Grok 4.7 xhigh (19.6%). Same id family as Grok 4.6. Sources: announcement + model card PDF. • Self-reported
19.6%
Harvey's Legal Agent Benchmark
Harvey's Legal Agent Benchmark; Grok 4.7 xhigh. Score reported by xAI. • Self-reported
19.6%
HealthBench Professional
HealthBench Professional; Grok 4.7 xhigh (56.7%). Sources: announcement + model card PDF. • Self-reported
56.7%
LAB-Bench Practical
LAB-Bench practical MCQ accuracy; Grok 4.7 high (76.8%). Distinct from labbench2. Sources: model card PDF. • Self-reported
76.8%
LatchBio Capabilities v1.0
LatchBio Capabilities v1.0 overall; Grok 4.7 xhigh (44.5%), equal-weight mean of 11 capability benches. Sources: model card PDF. • Self-reported
44.5%
ProtocolQA Open-Ended
ProtocolQA Open-Ended accuracy; Grok 4.7 high (70.4%). Distinct from MCQ ProtocolQA. Sources: model card PDF. • Self-reported
70.4%
SWE-Marathon
SWE-Marathon v1.1; Grok 4.7 high (46.0%). Sources: model card PDF. • Self-reported
46.0%
Terminal-Bench 4.0
Terminal-Bench 4.0; Grok 4.7 xhigh (38.0%) with Grok Build harness. Sources: announcement + model card PDF. • Self-reported
38.0%
Virology Capabilities Test
Virology Capabilities Test (VCT) accuracy; Grok 4.7 high (63.0%). Capability signal, not refusal. Sources: model card PDF. • Self-reported
63.0%
WMDP-Bio
WMDP-Bio accuracy; Grok 4.7 high (88.1%). Dual-use biology MCQ capability signal. Sources: model card PDF. • Self-reported
88.1%
WMDP-Chem
WMDP-Chem accuracy; Grok 4.7 high (84.9%). Dual-use chemistry MCQ capability signal. Sources: model card PDF. • Self-reported
84.9%
WMDP-Cyber
WMDP-Cyber accuracy; Grok 4.7 high (88.1%). Dual-use cyber MCQ capability signal. Sources: model card PDF. • Self-reported
88.1%

License & Metadata

License
proprietary
Announcement Date
September 21, 2026
Last Updated
September 22, 2026

Compare Grok 4.7

All comparisons

Similar Models

All Models

Recommendations are based on similarity of characteristics: developer organization, multimodality, parameter size, and benchmark performance. Choose a model to compare or go to the full catalog to browse all available AI models.