Grok 4.6 Ships the Week After helloai Dropped It
xAI ships Grok 4.6 today at the same $2/$6 as 4.5, five days after helloai dropped Grok. Arena votes are still too thin to reopen the slot.
xAI shipped Grok 4.6 today, five days after helloai dropped Grok 4.5 from the six-model frontier table. Official docs list grok-4.6 at the same $2 per million input tokens and $6 per million output that 4.5 charged, with the same 500K-token context window and a $4/$12 surcharge once a prompt crosses 200K. The company recommends it for both chat and code. It is live on the SpaceXAI API plus Cursor, Grok Build, OpenRouter, Vercel, and Cloudflare.
The vendor evals are why 4.6 exists. On xAI's launch table, Grok 4.6 High ties GPT-5.6 Sol Max at 61 on the Artificial Analysis Intelligence Index — one point behind Fable 5 Max at 62 — and lifts CursorBench v3.2 from 66.7 percent on 4.5 High to 69.9 percent. Terminal-Bench v3.0 jumps from 15.7 percent to 26 percent. Those are xAI's numbers, not independent reruns, and DeepSWE v1.1 still trails Sol (65.9 percent versus 73 percent). The pitch is longer agent trajectories and stronger first passes on visual, interactive projects, not a new price or context class.
Arena is the part that keeps 4.6 off helloai's board this week. The text-overall listing for grok-4.6-high sits at 1464 Elo with a plus-or-minus 12 interval and 2,448 votes, marked Preliminary. That is four points below the 1468 we last recorded for Grok 4.5, and it is a fresh debut. Admission requires Elo within 25 points of the tracked floor sustained for two weeks, not a launch-day sample. DeepSeek V4 Pro, which took the slot on August 7, is still the cheapest tracked API at $0.435/$0.87 with 1M context.
helloai is not saying 4.6 is weak. Same-provider upgrades usually replace the prior xAI row automatically; there is no xAI row to replace, because last week's six-model cut went to DeepSeek's price leadership. Re-admitting Grok now would mean dropping someone else — DeepSeek, Gemini 3.1 Pro, or Qwen3.8-Max — without two weeks of arena evidence that 4.6 holds a unique niche at $2/$6. Muse Spark 1.2 still undercuts that rate at $1.25/$4.25.
The first-week 2x included usage in Cursor and Grok Build is the right place to test the agent claims before any directory has to decide. If votes mature and 4.6 stays inside the Elo band while keeping a coding-agent story DeepSeek does not own, the next weekly pass can reopen the human-review question. Until then, /api/recommend keeps DeepSeek as Honest Daily Use and leaves Grok 4.6 as a same-price successor that has not yet earned its row back.