On September 21, 2026, SpaceXAI released Grok 4.7, its newest model for coding and knowledge work. The company describes it as the most capable model it has built for those tasks: it works longer on difficult problems, checks its own work more carefully, and carries what SpaceXAI calls its best-calibrated safeguards to date. The part most likely to get attention, though, is the price — Grok 4.7 is served at the same rate as its predecessor, about half what comparable frontier models charge.

Grok 4.7 succeeds Grok 4.6, which SpaceXAI announced in August. In a week that already saw Anthropic and OpenAI launch lower-priced models within hours of each other (Claude Opus 5.5 and GPT-6 Sol and Luna), SpaceXAI has joined the same argument: frontier capability is getting cheaper, and the race is now partly about who sells it for less.

A new base model, a longer training run#

Under the hood, the changes are straightforward. Grok 4.7 uses a new, larger base model than Grok 4.6, and it was trained with a longer reinforcement-learning run on a harder mix of tasks — one deliberately weighted toward problems that take many hours to complete rather than quick, single-turn answers. SpaceXAI says the result is a model that verifies its own work more reliably and holds longer context without losing the thread.

The model was also trained to natively understand the Grok Bot harness, which SpaceXAI says makes it stronger at conversational and general knowledge work. Grok Bot, introduced in August, is the company’s always-on agent setup — teams of agents with their own computer that work inside tools and apps around the clock. Grok 4.7 is also billed as better at creating documents and presentations.

SpaceX Starship on the launch pad with engines igniting
Starship on the pad during IFT-5. Photo by Steve Jurvetson, CC BY 2.0, via Wikimedia Commons.

The scorecard: competitive, not dominant#

SpaceXAI published a comparison table pitting Grok 4.7 against Grok 4.6, OpenAI’s GPT-5.6 Sol, and Anthropic’s Fable 5.1. One caveat first: Grok 4.7 ran at “xhigh” effort, Grok 4.6 at “high,” and the rivals at “max” — company-reported scores with uneven settings, so treat the gaps as fuzzy:

BenchmarkGrok 4.7Grok 4.6GPT-5.6 SolFable 5.1
CursorBench 4.0 (longer-running coding)46.3%40.4%41.7%51.8%
DeepSWE v1.1 (software engineering)71.0%*65.2%72.7%70.0%
Terminal-Bench 4.0 (multi-hour terminal work)37.6%20.3%37.3%57.9%
EEBench (electrical engineering)64.0%53.0%39.4%56.4%
AA Briefcase v1.1 (multi-hour office work)1,6571,5461,4871,678
Harvey Legal Agent Benchmark19.6%15.8%2.5%6.7%
HealthBench Professional (clinical reasoning)56.7%48.5%60.5%62.1%
GDPval (professional knowledge work)1,6951,605—1,735

All figures are company-reported. Grok 4.7 ran at xhigh effort, Grok 4.6 at high, and GPT-5.6 Sol and Fable 5.1 at max; the DeepSWE figure carries a high-effort footnote. A separate GDPval chart puts GPT-6 Astra at 1,542.

The honest read: Grok 4.7 is clearly better than its predecessor and roughly in the frontier pack — ahead on some evaluations, behind on others. SpaceXAI’s own framing is price-performance, and that is the right lens: not a new ceiling, but frontier-adjacent capability at a mid-tier price.

Price is the real argument#

Grok 4.7 starts at $2 per million input tokens and $6 per million output tokens — identical to Grok 4.6, and well under the $4/$20 listed for GPT-5.6 Sol and the $10/$50 for Fable 5.1. A fast variant doubles output speed at twice the price. It is available in Cursor, the Grok API, third-party harnesses, routers and cloud platforms — and free inside Grok Build, SpaceXAI’s own coding agent.

The timing is impossible to miss: three frontier-class launches in two days, all arguing intelligence is getting cheaper rather than smarter. For teams running long agentic sessions, where per-task cost compounds fast, that shift matters more than a few benchmark points.

  • Who this helps: teams running long agentic coding sessions, where per-task cost compounds fast.
  • Who should wait: security and life-science teams — the safeguard claims are company-reported, and independent testing will be the real verdict.
  • What to benchmark yourself: multi-hour tasks at default effort, where the scorecard’s asterisks do the most hiding.
Aerial view of a large AI datacenter campus at sunset
Compute is the price argument: hyperscale datacenter capacity like this powers the longer training runs. Photo by Chad Davis, CC BY 2.0, via Wikimedia Commons.

A new safeguard stack, tested on refusal#

The other half of the announcement is safety. SpaceXAI says Grok 4.7 was built with an entirely new safeguard stack and calls it the strongest model the company has tested on refusals and jailbreak resistance. In dual-use domains such as cybersecurity and biological work, the company says it leads on both helping with benign tasks and refusing dangerous ones, topping LatchBio’s biosafety benchmark at 62.4%.

On HackerBench v0.3, SpaceXAI’s own benchmark for risky and malicious cyber tasks, the company reports Grok 4.7 let only 3.3% of risky dual-use prompts through while rarely blocking legitimate security work. It has also started giving select cybersecurity partners invite-only access to the model’s red-team capabilities for defense research. The benchmark claims will need independent red teams to carry weight.

What to watch#

Three things will decide whether Grok 4.7 matters. First, independent evaluations: will the “works longer, checks its own work” claim survive third-party testing on multi-hour tasks? Second, distribution: Grok 4.6 reached GitHub Copilot, Amazon Bedrock, the Gemini Enterprise Agent Platform, and Microsoft Foundry within weeks — similar pickup would put 4.7 in front of far more developers. Third, the price war: if frontier-class coding keeps sliding toward $2 per million input tokens, the labs’ ability to profit from the models themselves keeps sliding with it.