SpaceXAI has pushed its Grok models back to the front of the AI pack, releasing Grok 4.6 on Wednesday with a claim that lands harder than the usual launch boast, that it now matches the best models from OpenAI and Anthropic while charging roughly half as much. On the independent Artificial Analysis Intelligence Index, a composite of nine benchmarks, SpaceXAI says Grok 4.6 scored 61, drawing level with OpenAI’s GPT-5.6 Sol and finishing a single point behind Anthropic’s Fable 5, a near tie at the frontier that its price then undercuts.
That price is the sharpest part of the case, because SpaceXAI set Grok 4.6 at 2 dollars per million input tokens and 6 dollars per million output tokens, with a faster variant at double the rate, which the company frames as about half the cost of comparable frontier models from its rivals. For developers who pay by the token, matching the leaders on quality while roughly halving the bill is the kind of trade that moves real usage, and SpaceXAI is sweetening it further by giving Cursor and Grok Build users double their included Grok 4.6 allowance during the first week, alongside access through Grok Bot, the SpaceXAI API, and third party platforms such as OpenRouter, Vercel, and Cloudflare.
The benchmarks tell a more textured story than the headline tie, because Grok 4.6 is at its strongest on tasks that resemble real professional work rather than isolated questions, leading both rivals on the GDPVal-AA v2 test with a score of 1753 and on the AA-Briefcase and Harvey LAB benchmarks, all of which measure the kind of multi step research, analysis, and document work that agents are increasingly asked to do. On pure coding the picture is more mixed, since Grok 4.6 edged out GPT-5.6 Sol on CursorBench with 69.9 percent to Sol’s 67.2, yet fell just short of Fable 5 at 70.5, and it trailed both models on the DeepSWE and Terminal-Bench coding evaluations.
SpaceXAI credits the gains to a longer supplemental training run than Grok 4.5 received, one that folded in curated engineering data and a wider round of reinforcement learning across agentic tasks spanning coding, knowledge work, and computer aided design. The company also says the model does more self testing on long tasks, checking its own work before it continues, which is a direct attempt to solve one of the hardest problems in agentic AI, the way small errors compound across a long chain of steps until the whole task drifts off course. That focus on long running agents and, as the company puts it, more ambitious interactive and visual work, is where SpaceXAI has chosen to compete hardest.
The market read the release as a genuine threat to the incumbents, since SpaceX stock climbed more than 6 percent after the announcement to trade near 142 dollars, and the move came days after Morgan Stanley laid out a 600 dollar bull case for the company, arguing that its AI business could grow into a broad platform with Grok as the intelligence layer underneath it. Elon Musk, never one to undersell, called the release a “banger” on X, and SpaceXAI shows no sign of slowing, with Grok 4.7, reported to be a 2.1 trillion parameter model, said to be only weeks away.
For all the noise around it, the substance is what makes this release notable, because a model that ties the best of OpenAI and Anthropic on a neutral benchmark while costing half as much changes the calculation for anyone choosing a frontier model, especially for the agent heavy work where Grok 4.6 looks strongest. The caveats are real, since it still trails on some coding tests and sits a point behind Fable 5 overall, and a benchmark tie is not the same as a lead. Even so, SpaceXAI has done the thing that matters most in this race, closing the quality gap while opening a price one, and it has promised to move again within weeks.

Your First 10 AI Skills
10 practical AI skills, copy-paste prompts and a 7-day plan to start using AI with confidence.
Download the guide →
