Blog Post

Grok 4.6

A few weeks ago, x.AI released Grok 4.6. I’ve recently started to use it, and I have not liked it too much.

Grok 4.6 is a good model. Extremely cheap and intelligent. However, compared to Grok 4.5, the token usage has increased. On the Artificial Analysis Intelligence Index (AAII), the average tokens per task for Grok 4.5 was 15k tokens. However, Grok 4.6 has an increased usage of 22k tokens, a roughly 50% increase in token utilization.

Notion image

This has affected real work. Grok 4.5 is noticeably faster than Grok 4.6 for tasks. Nevertheless, its usage is basically unlimited in Cursor, so using it is worth it.

My biggest concern with it is steering. If you interrupt the model while its running a task, it will go, do the task you have told it, and completely forget about the task before interruption. On the other hand, GLM-5.3-Flash, a model ~4.7x smaller than Grok 4.6 (Elon Musk has said that Grok 4.6 is 1.5T parameters, and GLM-5.3-Flash is 320B parameters), handles interruptions perfectly fine. While its token usage is quite egregious, its speed and pricing MORE than make up for it, hovering at around 60-70 tokens per second on providers and averaging $0.09 per task.

Notion image
Notion image
Notion image

However, at slower speeds, such as 43 tokens per seconds as measured by the AAII, can cause the time per task to spike to over 13 minutes, where both Grok 4.5 and 4.6 had a more than double improvement in time.

Notion image
Notion image
Notion image

I’m excited about Grok 4.7, which Elon Musk claims to be a 2.1T parameter model, and considering the fact that Kimi K3 is roughly double the parameter count over Grok 4.6, yet Grok 4.6 beats it out in evaluations (I personally have not tried Kimi K3, I will soon), I’m excited for what is to come with Grok 4.7.