Added to our archive on 1 October 2026. This report is dated to the day of the event.

SpaceXAI left the price of its new model exactly where it was: USD 2 per million input tokens and USD 6 per million output tokens. What changed with Grok 4.7, released on 21 September 2026, is how much work the model does for that money, and how many tokens it burns doing it.

Grok 4.7 is the company's "most capable model for coding and knowledge work", SpaceXAI wrote, and its first new flagship since SpaceX completed the purchase of the coding company Cursor in August. It went live on day one inside Cursor and in SpaceXAI's own coding tool, Grok Build, as well as through the Grok API. SpaceXAI is the former xAI: SpaceX took over Elon Musk's AI company in February, and SpaceX shares began trading on Nasdaq on 12 June. For developers the trade is easy to state and harder to price: better scores than Grok 4.6 at the same list price, but a model that thinks for longer.

  • What: Grok 4.7, SpaceXAI's new model for coding and long office tasks, built on a larger base model than Grok 4.6.
  • Price: USD 2 per million input tokens and USD 6 per million output tokens, unchanged; a fast variant costs twice as much for twice the output speed.
  • Why it matters: independent tests put it among the top four coding agents, but it uses more than twice the output tokens of Grok 4.6 per task.

What SpaceXAI changed

According to SpaceXAI's announcement, Grok 4.7 "uses a new, larger base model" than Grok 4.6 and was trained with a longer reinforcement learning run "weighted toward problems that take many hours to complete". The company says the model checks its own work more carefully, handles longer context better and was trained to work natively inside Grok Bot, SpaceXAI's agent product. It did not publish a parameter count. The context window stays at 500,000 tokens and cached input costs USD 0.50 per million tokens, both unchanged from Grok 4.6, Artificial Analysis reported.

On safety, SpaceXAI said Grok 4.7 has "an entirely new safeguard stack" and called it "the strongest model we've tested on refusals and jailbreak resistance". On HackerBench v0.3, the company's own test of risky cyber requests, it let through 3.3 percent of risky dual use prompts. Selected security partners get invite only access to its red team capabilities for defence research.

The vendor's scoreboard

SpaceXAI's comparison table shows big gains over Grok 4.6 and a mixed picture against rivals. The figures below are vendor reported, at the reasoning settings SpaceXAI chose (xHigh for Grok 4.7, Max for the two rivals). The table pits Grok 4.7 against OpenAI's GPT-5.6 Sol and Anthropic's Claude Fable 5.1. OpenAI's GPT-6 Astra appears only in a separate GDPval chart of professional tasks, where SpaceXAI put it at 1,542 Elo against 1,695 for Grok 4.7 and 1,735 for Fable 5.1.

Test (vendor reported)Grok 4.7Grok 4.6GPT-5.6 SolClaude Fable 5.1
CursorBench 4.0 (coding)46.3%40.4%41.7%51.8%
DeepSWE v1.1 (software engineering)71.0% (high effort)65.2%72.7%70.0%
Terminal-Bench 4.0 (multi hour terminal work)37.6%20.3%37.3%57.9%
EEBench (electrical engineering)64.0%53.0%39.4%56.4%
HealthBench Professional56.7%48.5%60.5%62.1%
Price listed by SpaceXAI, USD per million input / output tokens2 / 62 / 64 / 2010 / 50

Grok 4.7 leads on electrical engineering and on Harvey's legal agent test, where SpaceXAI reports 19.6 percent against 6.7 percent for Fable 5.1. It trails Fable 5.1 clearly on long terminal work and on CursorBench, the coding test built by Cursor itself.

What independent testers found

Artificial Analysis, which runs the same standardised tests across models, published its results on launch day. Grok 4.7 scored 46 on its Intelligence Index, two points above Grok 4.6. The gains were larger on long office work: 1,657 Elo on the firm's AA-Briefcase benchmark, 111 points more than Grok 4.6 and "just behind Claude Opus 5 and Claude Fable 5.1". Paired with Grok Build, it scored 56 on the firm's Coding Agent Index, nine points up, which ranks it fourth among models running in their own coding tools, behind Claude Fable 5.1, GPT-6 Astra and Claude Opus 5.

81,000
Average output tokens Grok 4.7 used per Intelligence Index task, against about 36,000 for Grok 4.6 at high effort and 27,000 for GPT-6 Astra (Artificial Analysis, 21 September 2026).

The catch is consumption. Grok 4.7 used more than twice the output tokens of Grok 4.6 per task and roughly three times as many as GPT-6 Astra. At an unchanged price per token, more than twice the tokens means a bigger bill for the same job, and a slower one: the firm measured an average of about 7.1 minutes per index task. Accuracy on its AA-Omniscience knowledge test was broadly flat, while the hallucination rate fell to 29 percent from 34 percent.

Why Cursor matters

Distribution is what sets this release apart. SpaceX agreed a deal with Cursor in April that gave it an option to buy the company for USD 60 billion, and in mid August Cursor announced that it was officially part of SpaceX, with "access to the largest fleet of GPUs in the world", TechCrunch reported. Grok 4.7 was available in Cursor on launch day, TestingCatalog noted.

One of Cursor's other suppliers has already reacted. On 29 August OpenAI said it would stop supplying models to Cursor on 12 November, explaining that it could not be confident SpaceX would use its technology within its terms of service, "based on our experience with Elon Musk's companies violating contracts", The Next Web reported. From that date OpenAI's models leave the editor, while SpaceXAI's own are built into it from day one.