SpaceXAI released Grok 4.7 on Monday 21 September 2026, billing it as the company’s most capable model for coding and knowledge work after a launch that slipped repeatedly since August. The headline claim is frontier price-performance: the model ships at $2 per million input tokens and $6 per million output — unchanged from Grok 4.6, and well below the $4/$20 that OpenAI’s GPT-5.6 Sol Max or the $10/$50 of Anthropic’s Fable 5.1 Max charge per SpaceXAI’s own comparison.
What the benchmarks show — and where it trails
The benchmarks are a mixed picture rather than a clean sweep. Grok 4.7 leads the comparison table on electrical engineering (64.0% on EEBench) and legal work (19.6% on the Harvey Legal Agent Benchmark, more than seven times GPT-5.6 Sol Max’s 2.5%), and it posts 71.0% on DeepSWE v1.1 at high effort — competitive with the field but still behind GPT-5.6 Sol Max’s 72.7%. On multi-hour office work (AA Briefcase v1.1), its 1,657 sits between Fable 5.1 Max’s 1,678 and GPT-5.6 Sol Max’s 1,487.
But on the marquee software engineering benchmark — CursorBench 4.0, which stresses longer-running coding tasks — Grok 4.7 scores 46.3%, ahead of GPT-5.6 Sol Max (41.7%) yet clearly behind Fable 5.1 Max’s 51.8%. Terminal-Bench 4.0 tells a similar story: 38.0% is a large jump from Grok 4.6’s 20.3%, but Fable 5.1 Max reaches 57.9%. HealthBench Professional lands mid-pack at 56.7% against GPT-5.6 Sol Max’s 60.5%. SpaceXAI’s framing — “at the frontier in price-performance” — is doing accurate work: the scores are strong for the price, not across-the-board leadership.
Third-party data fills in the texture. Artificial Analysis rates the xhigh variant at 46 on its Intelligence Index, inside the top 20 of the 655 models it tracks, with a 500K-token context window and text-and-image input. One caveat it flags: the model is unusually verbose, generating roughly 240 million tokens across index evaluations against a median of 92 million — which matters at the margin, because verbosity erodes some of the per-token price advantage on long agentic runs.
A longer reinforcement run, checking its own work
Under the hood, SpaceXAI says Grok 4.7 uses a larger base model than Grok 4.6, trained with a longer reinforcement-learning run on a harder mix of tasks weighted toward problems that take many hours. The company highlights self-verification — the model checking its own work — and longer-context management as the two behavioural upgrades. Grok 4.7 was also trained natively on the Grok Bot harness, which SpaceXAI says improves conversational and general knowledge work alongside the coding gains.
The release closes out an unusual delay cycle. Grok 4.6 and 4.7 had been flagged for August, and the slip reportedly came down to additional reinforcement-learning tuning. That is a different failure mode from the norm — most launch delays in this industry are about compute or evaluation gaps; this one was about the post-training run itself, the part of the pipeline labs now treat as the main lever for agentic reliability.
A new safeguard stack, with history in the background
SpaceXAI claims Grok 4.7 was built with an entirely new safeguard stack and is “the strongest model we’ve tested on refusals and jailbreak resistance,” topping LatchBio’s biosafety benchmark at 62.4% while showing the highest safety on its HackerBench v0.3 for malicious cyber tasks. The claim sits against a patchy record for the model family: a Grok CLI agent uploaded a user’s entire home directory to xAI’s servers in July, and a BBC investigation documented delusional and harmful outputs from Grok chatbot sessions earlier this year. Whether the new stack closes that gap is not something a launch post can settle — it will show up in incident rates over the next quarter.
The pricing message, though, is unambiguous. Every point release in the Grok 4 line has held the $2/$6 price while the score rises, and rivals have responded in kind — Alibaba’s Qwen pushed agentic pricing cuts of its own this month. For New Zealand developers and businesses paying US-dollar token bills, the race is now about capability per dollar rather than capability per se. Musk’s stated next target remains Grok 5, which he has pegged at 10% odds of AGI on a six-trillion-parameter base — meaning today’s release is a step on a roadmap, not a destination.