The measurable change

SpaceXAI launched Grok 4.7 on September 21 for coding and knowledge work. On the company’s published CursorBench 4.0 comparison, Grok 4.7 at xhigh effort scored 46.3%, versus 40.4% for Grok 4.6 at high effort. The effort settings differ, so this is a reported product comparison rather than a controlled like-for-like experiment.

The company also reports 71.0% on DeepSWE 1.1 at high effort and 37.6% on Terminal-Bench 4.0, both higher than the Grok 4.6 figures shown in its launch material. These tests target sustained software tasks rather than a short code-completion prompt.

What changed under the hood

SpaceXAI says Grok 4.7 uses a larger base model and a longer reinforcement-learning run focused on difficult, multi-hour problems. It also says the model has improved self-checking and longer-context management, and was trained to understand the Grok Bot harness natively.

List prices start at $2 per million input tokens and $6 per million output tokens, the same listed price as Grok 4.6. The model is available through Grok’s API and supported coding tools. A faster serving variant costs more, so teams should compare the end-to-end cost of a completed task rather than assume one advertised speed or price applies to every configuration.

Why this matters

Long jobs fail in ways short benchmarks can hide: an agent forgets an earlier constraint, edits the wrong file, or stops checking its own work. Better context and verification could help, but the visible benchmark figures still need to be tested against real repositories and human review standards.

The practical experiment is to give Grok 4.6 and 4.7 the same well-scoped issue, then compare merged results, time to completion, test quality and reviewer effort. The frontier is moving from producing code snippets toward carrying through a sequence of engineering decisions; this release is another test of whether that shift holds outside the lab.

Explore the original source ↗

Source published 2026-09-21. Coverage is based on the maker’s announcement and demonstration.