Five slips, then a Monday. Grok 4.7 went live on 21 September after Musk had moved the date at least five times since late July, and the migration is one string: grok-4.7 bills the same $2 in and $6 out per million as 4.6, keeps the 500K window, and keeps the 200K band where the whole request reprices at $4 and $12. So the price story is short. The benchmark story isn't, because xAI's own table runs Grok 4.7 at xhigh effort and Grok 4.6 at high, and every gain on the page is measured across that gap.
The short answer
Grok 4.7 shipped 21 September 2026 at Grok 4.6 prices, with a new larger base model and a longer RL run. xAI's table has it ahead of 4.6 on all seven rows, but the 4.7 column is labelled xhigh and the 4.6 column high, so part of the jump is reasoning budget, not model. Against other labs it beats GPT-5.6 Sol on five rows and Claude Fable 5.1 on three, and it's 20 points behind Fable on Terminal-Bench 4.0.
Same rate card, one new model string
Here's everything that changed on the invoice: nothing. The models table lists grok-4.7 at $2.00 per million input, $0.50 cached and $6.00 output for prompts under 200K tokens, and $4.00, $1.00 and $12.00 at or above it. Those are the grok-4.6 numbers to the cent, including the cache rate that quietly rose 67% in August. The window is still 500K. The rule is still that a request whose prompt reaches 200K is billed at the higher rate for all of its tokens, not the overflow, so the arithmetic we did for 4.6 carries over untouched. If your agent drifts past 200K mid run, the bill doubles, and nothing in the response tells you.
The launch post also mentions a fast variant, "twice the output speed at twice the price", which works out to $4 and $12 below the band. We couldn't find a model id for it in the docs table, and neither could the trackers we checked. Treat it as announced rather than orderable until an id appears.
Availability is real on day one: the xAI API, Grok Build, Cursor, and the usual routers and cloud platforms. The docs give a knowledge cutoff of May 2026 for grok-4.7. The blog post doesn't mention one.
To confirm the model is visible to your key before you touch a config:
curl -s https://api.x.ai/v1/models/grok-4.7 -H "Authorization: Bearer $XAI_API_KEY"
Seven rows up, all of them xhigh against high
The table on the launch page has seven benchmarks and four models. On every row Grok 4.7 beats Grok 4.6. Terminal-Bench 4.0 goes from 20.3% to 38.0%, EEBench from 53.0% to 64.0%, CursorBench 4.0 from 40.4% to 46.3%, DeepSWE v1.1 from 65.2% to 71.0%, AA Briefcase from 1,546 to 1,657, the Harvey legal agent benchmark from 15.8% to 19.6%, and HealthBench Professional from 48.5% to 56.7%. Real gains, if the columns were comparable.
They aren't, quite. The column header says Grok 4.7 xHigh and Grok 4.6 High. We wrote in August that xhigh arrived with 4.6, so xAI had that setting available for the older model and didn't use it here. Nothing published lets you separate the model improvement from the extra thinking. The one row with a footnote, DeepSWE, was run at high for 4.7. That row gained 5.8 points, the second smallest relative gain on the table behind AA Briefcase, which is at least consistent with the effort level doing some of the work elsewhere.
Against other labs the picture is more useful, because those columns at least aren't xAI's own previous model. Grok 4.7 beats GPT-5.6 Sol on five rows, most sharply on EEBench (64.0% against 39.4%) and the Harvey legal benchmark (19.6% against 2.5%). It beats Claude Fable 5.1 on three: DeepSWE by a point, EEBench by 7.6, Harvey by 12.9. It loses to Fable on CursorBench, AA Briefcase, HealthBench, and by 19.9 points on Terminal-Bench 4.0, the benchmark xAI itself files under multi hour terminal work. That's the row we'd look at first for agent use, and it's the widest loss on the page.
Two more numbers come from the press rather than the page. CNET reports a GDPval score of 1,695 for Grok 4.7 against 1,735 for Fable 5.1, and a parameter count of 2.1 trillion, up from 1.5 trillion in 4.6. The parameter figures trace back to Musk's posts, and the launch page says nothing about size, nor about the Starlink telemetry and SpaceX failure logs that coverage says went into training. Maybe both are true. They're not on the page, and after watching another lab's launch table get edited for a day earlier this month, we'd rather quote the page and label the rest.
What we'd actually do with it
If you're on grok-4.6 already, the swap costs nothing on the rate card and you should make it, then re run your own evals at the effort you actually pay for. xAI's headline numbers are at xhigh, and xhigh means more output tokens at $6 per million. The launch page doesn't publish token counts per task, so we can't tell you how much more. Artificial Analysis will, probably within the week, and that's the number we'd wait for before believing "same price" means "same bill".
If you're choosing between labs for agent work, the honest reading of xAI's own table is that 4.7 has closed the gap to GPT-5.6 Sol and is now trading rows with it, while Fable 5.1 keeps a clear lead on the long terminal tasks. The EEBench and legal wins are large on this harness. I'm just not sure how many readers here pick a model on electrical engineering questions. On the coding rows, the ones we'd use, it's a mid table result.
And keep the 200K line where it was. Nothing about 4.7 moves it, and on a long running agent that single threshold still changes your bill more than the model does. Musk has already sketched 4.8 and 4.9, then Grok 5 as a possible frontier leader, with no dates on any of them. Given the summer we just had, we wouldn't plan around the next one either.
Sources
Release date, the fast variant wording, the Grok Bot harness note, and the full benchmark table with its effort labels and footnote come from xAI's Grok 4.7 announcement of 21 September 2026, whose announcement card is the image at the top. Pricing for grok-4.7 and grok-4.6, the cached rates, the 200K band, the 500K window and the May 2026 knowledge cutoff are from the xAI models documentation. The model id, the absence of a separate fast id and the effort levels were cross checked against LLM Stats, which also notes none of the scores are independently verified yet. The parameter counts, the reported training data, the GDPval figures, the timeline of delays and Musk's roadmap remarks are from CNET's launch coverage, and we could not confirm them on any xAI page.
Frequently asked questions
How much does Grok 4.7 cost?
Two dollars per million input tokens, fifty cents per million cached input and six dollars per million output, for prompts under 200K tokens. At or above 200K the whole request is billed at four dollars input, one dollar cached and twelve output. Those are the xAI documentation rates on 22 September 2026, and they're identical to Grok 4.6.
Is Grok 4.7 better than Grok 4.6?
On xAI's table, yes, on all seven benchmarks. But that table runs 4.7 at xhigh effort and 4.6 at high, so some of the gap is extra reasoning rather than a better model. Run both at the same effort on your own tasks before you decide how much of the gain is real for you.
Does Grok 4.7 beat Claude Fable 5.1?
On three of seven rows in xAI's own table: DeepSWE v1.1 by one point, EEBench by 7.6 points and the Harvey legal benchmark by 12.9. It trails Fable 5.1 on CursorBench 4.0, AA Briefcase, HealthBench Professional and Terminal-Bench 4.0, where the gap is 19.9 points. No independent evaluation had been published when we wrote this.
What is the Grok 4.7 fast variant?
xAI's launch post says it serves a fast variant with twice the output speed at twice the price, so $4 input and $12 output below 200K. We found no model id for it in the xAI models table, so as of 22 September you can't select it by name in the API. Check the docs table before you budget for it.
Is Grok 4.6 being retired?
Nothing on the launch page or in the models documentation announces a retirement date for grok-4.6. Both ids are listed at the same price, so there's no cost reason to stay, but there's also no deadline forcing a move.






















