← all news

Grok 4.6 matches top models on cost by using far fewer agent turns

AI · · · source (artificialanalysis.ai)

xAI released Grok 4.6 about a month after 4.5, and the interesting part is not the headline score but how it gets there. In Artificial Analysis's independent testing, the model lands at 61 on their Intelligence Index, level with GPT-5.6 Sol and just behind Claude Opus 5 at 63 and Fable 5 at 62. On its own that would be a routine update. What separates it is efficiency: on long-horizon agent tasks, Grok 4.6 finished in about 53 turns and 0.5 billion input tokens on average, against roughly 103 turns and 2.0 billion tokens for Claude Opus 5. Fewer turns and fewer tokens for similar quality means lower cost per task, and the measured figure came out at $0.84, the same as Kimi K3.

Pricing is $2 per million input tokens and $6 per million output, with cached input at $0.5. The context window stays at 500k tokens, unchanged from 4.5. On task benchmarks it is competitive rather than dominant: 88.4% on Terminal-Bench v2.1, level with the leaders, and a GDPval-AA Elo of 1753, second only to Opus 5.

The pattern worth watching is that a model can reach the front on price by being terse and decisive in agent loops, not by topping a single benchmark.

Why it matters

If you run agents at scale, token count per task is your bill. A model that reaches similar quality in half the turns and a quarter of the input tokens can cut cost sharply, so test Grok 4.6 on your own workload before assuming the highest-scoring model is the cheapest to operate.

xAIModelsAgents