OpenAI turned on Ultrafast for GPT-6.1 Sol on 8 October 2026, and the number to remember isn't the speed. It's six. Every rate on the Sol price list gets multiplied by six: $12 per million input tokens, $60 per million output, against $2 and $10 on the standard tier. OpenAI's own announcement frames that as "1.2x the cost of Astra", which is true only if you compare it with Astra's standard price.
The short answer
Ultrafast on gpt-6.1-sol costs $12 in, $60 out, $0.60 per cached read and $15 per cache write, per million tokens. That's 6x the standard tier, 3x Fast mode and 1.2x GPT-6 Astra's standard price. You opt in per request with service_tier: "ultrafast". OpenAI promises up to 6x faster generation in the API, and the guide doesn't say how long contexts are billed.
What the price list says
Here's the Sol column on OpenAI's pricing page, per million tokens, for requests up to 272,000 input tokens. Standard: $2 input, $0.10 cached input, $2.50 cache write, $10 output. Fast: exactly double, so $4, $0.20, $5 and $20. Ultrafast: $12, $0.60, $15 and $60. Every line is six times standard, including cache reads and cache writes, so a prefix you cache stays six times dearer to read.
The "1.2x Astra" line needs unpacking. GPT-6 Astra's standard price is $10 and $50, and $12 and $60 is 1.2 times that. But Astra has its own Ultrafast tier at $60 in and $300 out, and next to that, Sol Ultrafast is a fifth of the price. So the honest comparison depends on what you were going to buy. If the alternative was Astra at normal speed, Sol Ultrafast costs a bit more and, per OpenAI, lands near Astra on quality while generating faster. If the alternative was Astra Ultrafast, it's a bargain. If the alternative was plain Sol, you're paying 6x for speed.
A concrete case, using only the list prices. An agent step that sends 50,000 uncached input tokens and gets 4,000 back costs $0.10 plus $0.04, so $0.14, on standard Sol. On Ultrafast it's $0.60 plus $0.24, so $0.84. Run a few thousand of those a day and the tier is a line item somebody will ask about.
What the speed claim covers
OpenAI says Ultrafast generates up to 8x faster in Codex, reaching 300 tokens per second, and up to 6x faster through the API. Both are "up to" figures. It also says the tier is meant for work like debugging an outage, agents driving apps and live experiences. Notice what's being sped up: the time between generated output tokens. Time to first token, queueing and your own tool calls aren't in that sentence.
That matters for agents. The guide strongly recommends WebSocket connections for agents making many tool calls in a row, because plain HTTP overhead can eat the latency advantage. I read that as a warning: if your loop spends most of its wall-clock time waiting on tools or on connection setup, a 6x token rate won't give you anything like a 6x faster run. I haven't benchmarked the tier, so treat any speedup number, mine or anyone's, as something to measure on your own traffic before you commit.
There's also a separate rate limit. Sol on Ultrafast gets 1M tokens per minute on the Build tier, 4M on Launch and 40M on Grow (the three usage tiers OpenAI introduced on 6 October). Astra gets 500K, 1M and 5M. Those limits sit apart from Standard and Fast, so a service that already runs near its Standard ceiling doesn't gain headroom by moving traffic across. Data residency covers the US and EU for Sol, and the US only for Astra.
The gaps in the documentation
Two things the guide leaves open. First, long context: Sol's standard price doubles input past 272K tokens, with $4 and $15 on the whole request. The Ultrafast guide says long-context behaviour isn't documented. Until a bill proves otherwise, I'd budget as if a long-context surcharge applies on top of the 6x, though I could be wrong and it might be simply not offered. Second, the guide doesn't compare Ultrafast with Fast mode, so what you get for the extra $8 per million output tokens over Fast is OpenAI's speed claim and nothing more detailed.
My take: Ultrafast is a sensible product for a narrow set of jobs, where a person is waiting and the output is long. For batch work, background agents and anything where a retry is cheaper than a premium, it's the wrong tier. The price tracker now carries the row, next to the earlier GPT-6.1 Sol cache cut and our piece on Astra's 272K cliff; the full table is in the LLM API pricing tracker.
Frequently asked questions
How do I turn Ultrafast on for GPT-6.1 Sol?
Set model to gpt-6.1-sol and service_tier to ultrafast on each request. It's per request, so you can send only latency-sensitive calls through it. OpenAI recommends WebSocket connections for agents making rapid tool calls.
How much does GPT-6.1 Sol Ultrafast cost?
$12 per million input tokens, $0.60 per million cached input tokens, $15 per million cache write tokens and $60 per million output tokens, for requests up to 272,000 input tokens. Standard is $2, $0.10, $2.50 and $10.
Is Ultrafast on Sol cheaper than Ultrafast on GPT-6 Astra?
Yes. Astra's Ultrafast tier is $60 in and $300 out, so Sol's $12 and $60 is a fifth of that. Sol's Ultrafast price is 1.2 times Astra's standard $10 and $50.
Does Ultrafast apply the long-context surcharge?
OpenAI's guide says long-context behaviour isn't documented for the tier. On standard Sol, input past 272K tokens bills $4 and $15 on the whole request. Don't assume Ultrafast is exempt.
Sources: OpenAI, Ultrafast mode guide, API pricing and API changelog; the OpenAI announcement of 8 October; independent report: Runtimewire. All read on 10 October 2026. The multiples and the 50,000-token example are our arithmetic from the list prices.

