xAI has released Grok 4.6, a retraining of Grok 4.5 focused on programming and long-running tasks. In independent comparisons it reaches the level of GPT-5.6 Sol for the first time. On ChatX it costs only about a third of that. This article shows the numbers, explains the quirk around thinking and says when the smaller Grok 4.3 remains the better choice.
Grok 4.5 against 4.6 in numbers
Grok 4.6 is not a new model generation. It is a longer training run on Grok 4.5 with data picked for reasoning, software development and technical subjects. The progress shows up in several tests:
| Test | Grok 4.5 | Grok 4.6 | What it measures |
|---|---|---|---|
| Artificial Analysis Intelligence Index | 56 | 61 | combined score from ten individual tests |
| DeepSWE | 54.0 % | 65.9 % | software tasks handled on its own |
| FrontierCode | 56.6 % | 61.3 % | hard programming problems |
| CursorBench | 66.7 % | 69.9 % | code changes in real projects |
| GDPval | 1,526 | 1,753 | professional knowledge work |
The biggest jump is in DeepSWE, the tasks a model works through on its own across many steps. That was exactly where Grok 4.5 was weak.
Where Grok 4.6 stands in the field
In the Intelligence Index, Grok 4.6 is level with GPT-5.6 Sol. It does not reach GPT-6 Astra, Claude Opus 5.5 or Fable 5.1, but for many tasks the gap is small enough to justify the price difference. Going by provider prices, output from Grok 4.6 costs around a third of Sol and around a third of Opus 5.5 as well. ChatX bills on the same prices, so for an answer of the same length that ratio also applies to your token deduction.
That makes Grok 4.6 interesting for anyone who has been using Sol for programming or technical research. For texts that need style and tone, Sol stays ahead. For code, web development and multi-step analysis, the difference is barely visible in our comparisons.
The model always thinks
Grok 4.6 is a pure reasoning model. It thinks before every answer, at low, medium or high depth, but always. On ChatX you therefore get exactly those three levels for this model. There is no “Off”, because xAI does not allow thinking to be switched off on 4.6. The thinking steps themselves stay invisible, yet they count as output tokens and show up in your usage.
For short questions that means a little more waiting and a few more tokens than with a model that answers straight away. For most tasks “Low” is enough. “High” pays off with tricky bugs or tasks that have several possible solutions.
Grok 4.6 or Grok 4.3
- Grok 4.3: answers without thinking steps, costs less than half of Grok 4.6 and is available to guests on ChatX. With one million tokens it has twice the context window. The right choice for translations, short texts, summaries and questions with one clear answer.
- Grok 4.6: for registered users, with image understanding and a selectable thinking depth. The right choice for programming, web development, technical questions and anything that takes several steps.
- The test: a task that Grok 4.3 solves differently on the second attempt than on the first is a task for Grok 4.6.
With Grok 4.6, xAI becomes a real alternative for work tasks for the first time, as long as you accept that the model always thinks first.
