SpaceXAI released Grok 4.7 on September 21, calling it its most capable model for coding and knowledge work, at the same $2 per million input tokens and $6 per million output tokens as Grok 4.6, per MarkTechPost. Prompts above 200,000 tokens cost $4 and $12, and the context window tops out at 500,000 tokens. Coding scores rose sharply: Terminal-Bench 4.0 to 38.0% from 20.3%, CursorBench 4.0 to 46.3% from 40.4%, and DeepSWE v1.1 to 71.0% from 65.2%, per VentureBeat.
Artificial Analysis measured Grok 4.7 at its highest reasoning setting using about 81,000 output tokens per task, 125% more than Grok 4.6 and 196% more than GPT-6 Astra. Output tokens are the expensive side of the bill, and a model that reasons longer pays for it there. On the Intelligence Index that comes to $3.74 per task for Grok 4.7 against $1.99 for GPT-5.6 Sol at its maximum setting, even though Sol's list price is higher. The per-token price held. Per finished task, Grok 4.7 costs nearly twice as much as a model with a higher rate card.
Longer reinforcement learning runs aimed at multi-hour tasks, a larger base model, and more self-checking all push output length up, and SpaceXAI kept the rate card flat while the model got wordier. Rate cards now mislead. Measure cost per completed task on your own workload instead.
Bottom Line
Grok 4.7 got better at code and more expensive to finish a task, at an unchanged rate card. Benchmark any model on cost per completed task with your own prompts before switching, because reasoning length now moves the bill more than the list price does.