|

SpaceXAI Releases Grok 4.6: A 500K-Context Frontier Model Tuned for Long-Running Agents, Coding, and Knowledge Work

SpaceXAI simply launched Grok 4.6. The launch is a post-training improve over Grok 4.5 quite than a bigger base mannequin. SpaceXAI held the inspiration fixed and spent the development on an extended supplemental coaching run, regenerated supervised fine-tuning trajectories, and reinforcement studying in agentic environments. Agents that keep on a job throughout many steps with out drifting. Grok 4.6 scores 61 on the Artificial Analysis Intelligence Index, up 5 factors from Grok 4.5 and tied with GPT-5.6 Sol Max. The mannequin takes 500,000 context tokens, is dwell right this moment in Cursor and Grok Build, and provides a brand new xhigh reasoning-effort stage above the ladder Grok 4.5 shipped with.

Is it deployable?

Yes, in manufacturing, with a bounded set of workloads. The mannequin is mostly accessible by way of the xAI API as grok-4.6, is the default mannequin in Grok Build, ships in Cursor on all plans, and is routable by way of OpenRouter, Vercel, and Cloudflare. There is not any open-weights launch and no self-hosting path, so air-gapped deployments are out.

  • Company stage: Seed-stage groups and indie builders can undertake it instantly, since Cursor and Grok Build want no harness work. Mid-market engineering orgs are the strongest match: API-only integration, mTLS authentication, batch and precedence processing are documented. Regulated enterprises ought to stage a pilot first — the seller’s model historical past is a dwell procurement query in a number of shopping for committees.
  • Industries: Software and developer tooling, semiconductor and kernel engineering, {hardware} and CAD-adjacent design, monetary analysis, and authorized evaluation. The coaching combine explicitly focused a number of of those.
  • Applications: Repository-wide refactors, migration brokers, research-and-synthesis pipelines over 500K-token corpora, first-pass utility scaffolding from a product temporary, GPU kernel optimization, and document-heavy information work.

What really modified

Grok 4.6 will not be a bigger base mannequin. SpaceXAI describes an extended supplemental coaching run than Grok 4.5 obtained, utilizing curated model-generated knowledge for reasoning and superior technical ideas, high-quality engineering knowledge, and an improved optimizer and coaching recipe.

Grok 4.5 was then used to regenerate supervised fine-tuning trajectories throughout reasoning-effort ranges, agent harnesses, and domains spanning STEM, software program engineering, and information work, with problematic traces filtered by model-based checks. Reinforcement studying adopted in agentic environments masking information work, basic coding, internet growth, computer-aided design, and kernel optimization.

The behavioral perception is a crucial one: on longer trajectories, SpaceXAI reviews extra self-testing and verification, with the mannequin checking its personal work earlier than shifting on. That is a vendor commentary from inside testing, not an independently measured outcome.

The mannequin takes 500,000 context tokens, accepts textual content and picture enter with text-only output, has no acknowledged textual content output restrict, and carries a February 1, 2026 information cutoff. reasoning_effort now helps low, medium, excessive (default), and a brand new xhigh stage. SpaceXAI didn’t publish a parameter rely for Grok 4.6.

Benchmarks: learn the losses first

On xAI’s launch desk, Grok 4.6 (High) scores 61 on the Artificial Analysis Intelligence Index, up from 56 for Grok 4.5 and tied with GPT-5.6 Sol Max. It leads the desk on GDPval-AA v2 (1753 Elo, versus 1526 for Grok 4.5), AA-Briefcase (1577, versus 1313), and Harvey LAB.

It trails on the coding rows that matter most to engineering groups. DeepSWE v1.1 lands at 65.9%, up 11.9 factors generationally however behind GPT-5.6 Sol Max at 73%. Terminal-Bench v3.0 reaches 26%, practically double Grok 4.5’s 15.7% and nonetheless final of the 4 listed fashions. CursorBench v3.2 is 69.9%, FrontierCode v1.1 Extended is 61.3%, and APEX-Agents is 57.5%.

Two issues to notice whereas evaluating. First, the desk’s bolded wins on GDPval-AA v2 and AA-Briefcase sit inside Artificial Analysis‘ printed confidence intervals — they’re statistical ties, not leads. Second, the comparability set excludes Anthropic’s Claude Opus 5, which at the moment tops that index. The disclosed losses are the extra dependable sign.

Pricing and entry

Per the release notes, Grok 4.6 payments $2 / $0.50 / $6 per 1M tokens (enter / cached enter / output) beneath 200K immediate tokens, and $4 / $1 / $12 above that threshold. The launch web page additionally references a quicker variant at double the value, with no separate mannequin ID printed. Grok Build and Cursor are providing 2× included utilization for the primary week.

Teams ought to set a prompt_cache_key (or the x-grok-conv-id header on Chat Completions). Without it, requests scatter throughout servers and cache hits turn out to be unreliable, so full enter value applies.

Interactive explainer