A futuristic data centre corridor with rows of glowing server racks under warm golden lighting
News

Grok 4.6 Joins the Frontier at Half the Price of Claude and GPT

Grok 4.6 ties GPT-5.6 Sol on intelligence at $2/$6 tokens — less than half what Claude Opus 5 or GPT-5.6 Sol charge. The real story is agent efficiency: 53 turns vs Claude's 103.

SpaceXAIGrokAI modelsAI benchmarksagentic AI

SpaceXAI released Grok 4.6 on August 12, and the model lands at 61 on the Artificial Analysis Intelligence Index — in line with OpenAI’s GPT-5.6 Sol Max and behind only Anthropic’s Claude Opus 5 (63) and Claude Fable 5 (62). That alone would be routine in a year where the frontier gets reshuffled weekly. What makes Grok 4.6 worth paying attention to is the price tag.

What is the Artificial Analysis Intelligence Index? It’s a third-party benchmark that aggregates performance across reasoning, coding, agentic tasks, and knowledge work into a single score. Think of it as a composite rating for AI models — the higher the number, the broader the competence. A score of 61 puts Grok 4.6 in the top tier alongside models from OpenAI and Anthropic.

Headline pricing stays at $2 per million input tokens and $6 per million output — unchanged from Grok 4.5. Claude Opus 5 charges $5/$25. GPT-5.6 Sol charges $5/$30. Grok 4.6 delivers effectively the same intelligence score as GPT-5.6 Sol at less than half the cost.

The agentic results are where it gets interesting

Grok 4.6’s strongest showing is on agentic work, not static reasoning. On GDPval-AA v2, the benchmark for real-world agentic knowledge work, it scores an Elo of 1,753 — behind only Claude Opus 5 and statistically tied with Claude Fable 5 and Qwen3.8 Max. It hits 50.7% on 𝜏³-Banking (multi-turn customer service with tool use) and 88.4% on Terminal-Bench v2.1.

The efficiency profile is the part that should worry Anthropic and OpenAI. On AA-Briefcase, the long-horizon knowledge work benchmark, Grok 4.6 completes tasks in roughly 53 turns and 0.5 billion input tokens on average. Claude Opus 5 (max) needs about 103 turns and 2 billion input tokens for the same work. That is half the turns and a quarter of the input tokens — a cost advantage that compounds well beyond the per-token price gap.

VentureBeat’s analysis notes this matters because enterprise AI deployments are shifting from isolated prompt-and-response interactions toward agents expected to maintain state, operate tools, modify code, and recover from problems across longer execution paths. A model that reaches a comparable answer in half the turns has a structural cost advantage.

Not a clean sweep

Grok 4.6 does not dominate every category. On Terminal-Bench v3.0, it scores 26% — up sharply from Grok 4.5’s 15.7%, but GPT-5.6 Sol Max and Fable 5 Max both reach 34%. On DeepSWE v1.1, GPT-5.6 Sol Max leads at 73% versus Grok’s 65.9%. The evidence supports a substantial upgrade over the previous generation more clearly than it supports across-the-board superiority over rival frontier models.

SpaceXAI itself acknowledges a methodological caveat: third-party scores in its comparison table use the best self-reported or publicly available results, so the comparison should not be read as a perfectly controlled four-model evaluation. That is an honest disclosure, and it tempers the headline.

The long-context pricing trap

There is a catch in the pricing that enterprises should read carefully. Grok 4.6 supports a 500,000-token context window, but prompts below 200,000 tokens get the headline $2/$6 rates. Once a prompt crosses 200,000 tokens, the rate jumps to $4/$12 — and the higher pricing applies to all tokens in that request, not just the overflow. For agents that accumulate context over long sessions, which is precisely the use case SpaceXAI is targeting, the cost could climb steeply.

Artificial Analysis places Grok 4.6 at $0.84 per task on its cost-per-task metric — on the Pareto frontier, but actually less economical than some cheaper models like OpenAI’s GPT-5.6 Luna, z.ai’s GLM-5.2, and Meta’s Muse Spark 1.2. The headline token price is not the whole story.

The brand problem

The Grok name carries baggage that no benchmark can wash away. As VentureBeat details, past incidents include antisemitic outputs in 2025, politically skewed responses about South Africa, implausibly flattering assessments of Elon Musk, and an ongoing Ofcom investigation into the Grok account’s image-generation capabilities. The European Commission has a separate Digital Services Act investigation open.

None of that establishes that Grok 4.6 repeats those behaviours. But enterprise procurement teams at banks, government agencies, and healthcare providers do not evaluate models in isolation from vendor history. SpaceXAI acquired xAI in February 2026 and rebranded, but the consumer-facing brand is still Grok. For organisations that cannot afford their AI supplier to become a brand-safety event, that is a real friction point.

Distribution as competitive moat

Grok 4.6 is available today in Grok Build (SpaceXAI’s answer to Claude Code), inside Cursor (which SpaceXAI recently acquired), and through partners including OpenRouter, Vercel, and Cloudflare. SpaceXAI is offering double the included usage during the first week. That distribution matters — developers can drop the model into existing coding-agent environments without building a new harness.

This builds on the earlier Grok Build release and the Grok 4.5 coding model launch, both of which signalled SpaceXAI’s push into the agentic coding space. The trajectory is clear: cheaper frontier intelligence, optimised for long-running agents, distributed through the tools developers already use.

What this means for the market

The frontier model market is splitting in two directions. Anthropic and OpenAI push higher intelligence at premium prices — Claude Opus 5 at $5/$25, Fable 5 at $10/$50. SpaceXAI is betting that the practical sweet spot is frontier-adjacent intelligence at commodity pricing. If Grok 4.6’s turn efficiency holds in production, that bet looks reasonable.

For NZ companies evaluating AI agents, the pricing compression is unambiguously good news. The gap between frontier-tier and mid-tier models has narrowed to the point where the question is no longer “can I afford frontier intelligence?” but “which frontier model fits my workflow and risk profile?” Grok 4.6’s brand baggage may rule it out for regulated industries, but for internal coding and knowledge work, the economics are hard to ignore.

📰 Sources

Sources: Artificial Analysis, VentureBeat, SpaceXAI