
SpaceXAI has released Grok 4.7, its latest artificial intelligence model for coding, agentic tasks and knowledge work, with improvements aimed at longer and more complex tasks. The model retains the same headline API pricing as Grok 4.6, starting at $2 per million input tokens and $6 per million output tokens. However, independent testing indicates that Grok 4.7 can consume significantly more tokens while completing tasks, raising questions about its actual cost for developers and enterprises.
Grok 4.7 was released on September 21 and is available through the xAI API, Cursor, Grok Build and other developer platforms. SpaceXAI describes it as its most capable model for coding and knowledge work, saying it was designed to work for longer on difficult tasks, verify its own work more carefully and manage longer contexts.
The model uses a larger base model than Grok 4.6 and was trained through a longer reinforcement-learning process focused on tasks that can take several hours to complete. SpaceXAI said the training placed greater weight on difficult problems requiring extended reasoning and tool use.
Grok 4.7 supports a 500,000-token context window, along with text and image inputs and text output. Its reasoning modes include low, medium, high and xhigh, with high set as the default. The model is also available through the public xAI API under the model name grok-4.7.
On pricing, Grok 4.7 starts at $2 per million input tokens and $6 per million output tokens for prompts below 200,000 tokens. Cached input tokens are priced at $0.50 per million tokens. For prompts above 200,000 tokens, the input and output rates increase to $4 and $12 per million tokens respectively.
The unchanged headline pricing is one of the main elements of the launch. However, VentureBeat reported that the amount of tokens consumed by Grok 4.7 can materially affect the final cost of completing a task. A model with a lower price per token can still generate a larger bill if it requires substantially more reasoning and output tokens to reach a result.
Independent testing reported by Artificial Analysis showed a notable increase in token usage. According to the testing cited in the reporting, Grok 4.7 at xhigh reasoning used around 81,000 output tokens per Intelligence Index task, compared with approximately 38,000 for Grok 4.6 at the same reasoning level. This means the unchanged per-token rate does not necessarily translate into an unchanged cost per completed task.
The same testing also showed improvements in coding-agent performance. Artificial Analysis reported that Grok Build with Grok 4.7 achieved a higher Coding Agent Index score than the previous Grok 4.6 configuration. The index combines benchmarks covering software implementation, terminal-based work and technical repository questions.
SpaceXAI’s own testing also reports improvements across several coding and professional-work benchmarks. Grok 4.7 scored 46.3% on CursorBench 4.0, compared with 40.4% for Grok 4.6. On DeepSWE v1.1, Grok 4.7 recorded 71% at high effort, compared with 65.2% for Grok 4.6. On Terminal-Bench 4.0, it recorded 38%, compared with 20.3% for its predecessor.
The model is therefore positioned around longer-running software development and knowledge-work tasks rather than simply delivering faster responses to short prompts. SpaceXAI said Grok 4.7 has been trained to better verify its own work and manage extended context, capabilities that are particularly relevant to coding agents working across large repositories or multi-step assignments.
Grok 4.7 is also being made available through coding environments. SpaceXAI said the model is available in Cursor and Grok Build, as well as through the Grok API, third-party coding harnesses, model routers and cloud platforms. A faster version is available through Cursor and Grok Build at twice the standard token rates.
For businesses evaluating AI coding tools, the pricing structure highlights an important distinction between token price and task cost. The amount a company ultimately spends can depend on how many reasoning and output tokens a model uses, how many attempts are required, whether tools are called repeatedly and how long an agent works before completing a task.
Grok 4.7 therefore enters the coding market with two simultaneous developments: measurable improvements over Grok 4.6 on several long-running coding benchmarks and unchanged headline token prices, alongside independent evidence that its greater token consumption can increase the cost of individual workloads.
The model’s actual economic impact will consequently depend on the type of work being performed. For tasks where its additional reasoning and coding capability reduces failures or manual intervention, higher token usage may have a different cost implication than for routine tasks that could be completed with fewer tokens. The available benchmark results provide an early indication of this trade-off, but actual costs will vary by workload and usage pattern.




