Grok 4.6 places SpaceXAI back at the frontier

Ronni Holmvig Strøm · 2026-08-16

SpaceXAI shipped Grok 4.6 on August 12, roughly a month after Grok 4.5, and it is back on the intelligence frontier. Artificial Analysis scores it at 61 on its Intelligence Index, level with GPT-5.6 Sol and behind only the Claude Opus 5 and Fable 5 tiers. The interesting thing is what it costs to

SpaceXAI shipped Grok 4.6 on August 12, roughly a month after Grok 4.5, and it is back on the intelligence frontier. Artificial Analysis scores it at 61 on its Intelligence Index, level with GPT-5.6 Sol and behind only the Claude Opus 5 and Fable 5 tiers. The interesting thing is what it costs to run, and why.

The Price Did Not Move

Holding pricing flat across a generation is unusual at the frontier. Intelligence gains almost always arrive with a bigger invoice. Grok 4.6 posts a 5-point Intelligence Index gain over Grok 4.5, and another 23 over Grok 4.3 earlier in the year, while its headline pricing stays exactly where 4.5 left it: $2 per million input tokens, $6 per million output. The models scoring within two points of it, Claude Opus 5 at $5/$25 and GPT-5.6 Sol at $5/$30, cost several times more on the dimension that dominates reasoning-heavy work, which is output tokens.

That gap is easy to read as a discount and harder to read as what it is, which is an engineering claim. A vendor can only hold price flat through a capability jump if the model has gotten cheaper to serve at the same time it got smarter. Grok 4.6's measured cost per task lands at $0.84, the same as Kimi K3 at slightly lower intelligence, which puts it on the intelligence-versus-cost Pareto frontier for every agentic evaluation Artificial Analysis runs. The number worth chasing is the mechanism underneath that $0.84.

It Finishes in Fewer Turns

The mechanism shows up cleanly on long-horizon work. On AA-Briefcase, Artificial Analysis's private benchmark of extended agentic knowledge tasks, Grok 4.6 resolves a task in roughly 53 turns and half a billion input tokens on average. Claude Opus 5, running at maximum, takes roughly 103 turns and two billion input tokens to reach a comparable answer. Half the turns. A quarter of the input.

This is the part that matters more than the ELO. Long agentic loops accumulate context with every turn, and input tokens are the tax you pay on all of it, over and over, as the transcript grows. A model that reaches the same place in half the turns is not merely faster. It is spending a fraction of the compute to get there, and that advantage compounds well past what the per-token price would suggest. For most of the past two years we watched the field optimize tokens per second, which was the right metric when a model answered a question and stopped. The metric that governs an agent running for a day is tokens per task, and Grok 4.6 is the clearest public case yet of a frontier model winning on it.

What makes this credible is that a third party measured it. The efficiency is not a founder's spec-sheet claim. It is the observed turn count and token count from an outside benchmark, and it corroborates the 1753 GDPval-AA v2 Elo that SpaceXAI cited at launch, a score Artificial Analysis independently placed behind only Claude Opus 5. The strong results and the cheap results are the same result, seen from two sides.

The Bet Grok 4.7 Is Making

Grok 4.6 also tells you what to watch for next. SpaceXAI has been open that Grok 4.7 will be a larger model, a fresh pre-train rather than another supplemental run on the current base, and that it will be "slightly slower to serve" while pushing token efficiency further still. On a spec sheet, slower-to-serve reads as a regression. Framed against tokens per task, it is the correct trade: a larger model that thinks in fewer, denser steps can finish a job for less total compute even when each token takes longer to produce.

Whether Grok 4.7 holds that efficiency at greater scale is exactly the question 4.6 has now made worth asking, because 4.6 turned the argument from a claim into a measurement. The frontier is crowded with models that are a point or two smarter than each other. The room with more space in it is the one Grok 4.6 just walked into: the same intelligence, finished in fewer turns.