Source: xAI — 2026-08-12
Summary
xAI released Grok 4.6, built for long-running agentic work — multi-step research, operating across an entire codebase, iterating a product idea into a polished build — rather than single-turn chat. It ships with a 500,000-token context window and launched the same day into Cursor, xAI's own "Grok Build" tool, the xAI API, OpenRouter, Vercel, and Cloudflare, priced at $2 per million input tokens and $6 per million output tokens. On the Artificial Analysis Intelligence Index, a 9-benchmark composite, it scores 61, tying OpenAI's GPT-5.6 Sol, trailing Anthropic's Claude Fable 5 Max at 62, and beating the prior Grok 4.5 High's score of 56.
Key Takeaways
- The 500,000-token context window is the headline spec — large enough to hold an entire codebase or a long multi-step research trail without losing earlier context.
- Day-one distribution across six surfaces (Cursor, Grok Build, xAI API, OpenRouter, Vercel, Cloudflare) signals xAI is prioritizing developer reach over a single flagship app.
- At $2/$6 per million input/output tokens, it undercuts typical frontier pricing while scoring competitively on the composite intelligence index.
- The Artificial Analysis Intelligence Index places Grok 4.6 in a tight cluster with GPT-5.6 Sol (both at 61), just behind Claude Fable 5 Max (62) — the frontier gap between top labs has narrowed to a single point on this measure.
- The jump from Grok 4.5 High (56) to Grok 4.6 (61) is a 5-point gain on the index, xAI's largest generational move in this specific benchmark to date.
Reel Script
Hook (17s)
xAI just released a model built to work for an hour straight without losing the thread — and it landed inside Cursor, Vercel, and four other tools on the exact same day. That's the part worth paying attention to.
Core Concept (85s)
Most "new model" announcements are about being smarter on a single question. Grok 4.6 is pitched differently — it's built for agentic work, meaning tasks where the model has to take many steps in a row, remember what it did three steps ago, and keep going without a human re-explaining the goal each time. Think multi-step research, working across an entire codebase file by file, or taking a rough product idea all the way to a working build. The thing that makes long agentic runs possible is context window size — literally how much text the model can hold in its "working memory" at once, measured in tokens, which are roughly word-sized chunks. Grok 4.6 ships with 500,000 tokens of context, which is enough to hold a genuinely large codebase or a long chain of research steps without the model forgetting earlier decisions. And crucially, xAI didn't just publish a blog post — they shipped it live, same day, into the actual tools developers already use: Cursor, their own Grok Build tool, the API, OpenRouter, Vercel, and Cloudflare.
Hands-On (95s)
Let's put it next to the competition using one number: the Artificial Analysis Intelligence Index, a composite score built from nine different benchmarks blended together, so it's less "gameable" than any single test. Grok 4.6 lands at 61 on that index. GPT-5.6 Sol from OpenAI also scores 61 — dead tie. Claude Fable 5 Max from Anthropic scores 62, just one point ahead — effectively the current frontier leader, but barely. And the prior Grok model, 4.5 High, scored 56, so this release is a real 5-point jump generation over generation. On pricing, Grok 4.6 runs $2 per million input tokens and $6 per million output tokens, which is on the cheaper end for a model performing at this tier.
Takeaway (23s)
A three-way tie at the top of the intelligence index means model choice increasingly comes down to context window, tooling integration, and price — not raw smarts. If you're building agents that run long tasks across a big codebase, Grok 4.6's context window and day-one Cursor support make it worth a real trial this week.