• Latest
  • Trending
  • All
Cerebras launch artwork for the Ultrafast tier, dark background with orange and blue light trails and the title Accelerating GPT-5.6 Sol Ultrafast with OpenAI.

GPT-5.6 Sol Ultrafast hits 750 tokens a second, with no price

3 September 2026
The official xAI announcement card for Grok 4.7, white type on a dark grey and navy gradient.

Grok 4.7 keeps $2 and $6, and its gains over 4.6 are xhigh versus high

22 September 2026
Answer card stating that Qwen-Image-2.1, released on 20 September 2026, ships open weights with a 7 billion parameter diffusion transformer, a Qwen3-VL 8B text encoder and an RGBA VAE totalling about 33 gigabytes in BF16, under the Qwen Research License that limits use to research or evaluation and requires a separate commercial licence, unlike the Apache 2.0 licence of Qwen-Image 1.0.

Qwen-Image-2.1 brings the weights back, but not the Apache licence

21 September 2026
Answer card stating that Ternary Bonsai 2 27B, released by PrismML on 17 September 2026 under Apache 2.0, packs Qwen3.8 27B into 5.95 gigabytes at 1.72 bits per weight, keeps 98.2 percent of the 14-benchmark average, about 75 percent on SWE-bench Verified and Terminal-Bench 2.1, and needs PrismML's llama.cpp fork to run.

Does Bonsai 2 27B really keep 98% of Qwen3.8 in 5.95 GB?

20 September 2026
Answer card stating that Jev 1.13 from TypeSafe AI is a decision model in early access since 15 September 2026 that returns typed probabilities instead of text, priced at 42 dollars per billion input tokens with output tokens free, answering in 70 to 500 milliseconds, with a 64K token request budget, text input only, and a documented list of things it does badly, including counting and dates.

Jev 1.13 bills $42 a billion tokens, and it can’t count

19 September 2026
Answer card stating that Qwen3.8-Omni-Flash launched on 17 September 2026 as an API only model on Alibaba Cloud Model Studio, taking text, images, audio and video in a 1M token context and returning text only, priced at 0.15 dollars per million input tokens for every modality and 0.47 dollars per million output tokens in the international regions, with no open weights published and the Qwen-Live Harness GitHub repository returning 404.

Qwen3.8-Omni-Flash bills audio at $0.15 and ships no weights

18 September 2026
Answer card stating that on 15 September 2026 AWS said it is unable to restore access to resources and data hosted exclusively in the Middle East Bahrain region me-south-1 and in the mec1-az2 zone of the UAE region, because the damage spanned multiple Availability Zones and exceeded what multi-AZ services are designed to withstand.

AWS can’t restore me-south-1, six months after the drone strikes

17 September 2026
Answer card stating that Google released Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking on 15 September 2026 at 3 dollars per million audio input tokens and 12 dollars out, that the thinking model requires asynchronous tools, and that Artificial Analysis scores it 82.6 on its Speech to Speech Quality Index.

Gemini 3.8 Live Extended Thinking rejects any tool that blocks

16 September 2026
Answer card summarising the Atria Dawn Preview release: 744B GLM-5.2 base, MIT licence, 1.5 TB BF16 and 756 GB FP8 checkpoints, 256K context, top on five of sixteen benchmark rows and trailing on SWE-bench Pro.

Atria Dawn Preview is 744B under MIT, and the BF16 weighs 1.5 TB

15 September 2026
Answer card stating that OpenAI released the Agents API in public beta on 10 September 2026 with no separate fee, billed through model tokens, tool calls and hosted sandbox time, with a choice of OpenAI hosted, self hosted or partner sandboxes, US only data residency and no Zero Data Retention support.

OpenAI’s Agents API has no fee, no ZDR and a one hour sandbox clock

14 September 2026
Answer card: Sakana Fugu Max at $2 and $6 per million tokens, Fugu Ultra v2 unchanged at $5 and $30, and Sakana saying Ultra v2 scores without Fable 5 or GPT-6 Astra in its pool.

Fugu Max costs $2 and $6 while Fugu Ultra v2 runs without Fable 5

13 September 2026
Answer card stating that DeepSeek released DeepSeek-V4.1-Flash on 10 September 2026 as a 552 billion parameter mixture of experts model with a new causal encoder decoder architecture that activates 8 billion parameters on input and 16 billion on output, with native vision, a one million token context and MIT licensed weights, that the API model name is now deepseek-flash at 0.15 dollars per million input tokens and 0.60 dollars per million output tokens off peak, and that DeepSeek announced V4 Pro would be routed to V4.1-Flash from 14 September and reversed that on 11 September.

DeepSeek V4.1-Flash arrived, and the V4 Pro retirement lasted a day

12 September 2026
Answer card stating that Cognition released SWE-2 on 10 September 2026, a coding model post-trained from Kimi K3, scoring 50.0 percent on FrontierCode 1.1 Main against 50.9 percent for Claude Fable 5.1 and 27.3 percent on Terminal-Bench 4 against 55.8 percent, available only inside Devin.

SWE-2 trails Fable 5.1 by one point, and by 28 on Terminal-Bench 4

11 September 2026
  • About
  • Contact
  • Privacy
  • Legal
Wednesday, September 23, 2026
  • Login
Packet Nebula
  • Home
  • Articles
    • Security
    • Network
    • Dev
    • Sysadmin
    • SEO
    • Email & DNS
  • Tools
    • Network tools: free, fast, no signup
    • Security tools: free, fast, no signup
    • Developer tools: free, fast, no signup
    • Sysadmin tools: free, fast, no signup
    • SEO tools: free, fast, no signup
    • Email & DNS tools: free, fast, no signup
  • Download
  • About
No Result
View All Result
Packet Nebula
No Result
View All Result
Home Dev

GPT-5.6 Sol Ultrafast hits 750 tokens a second, with no price

by stephane
3 September 2026
in Dev
0
Cerebras launch artwork for the Ultrafast tier, dark background with orange and blue light trails and the title Accelerating GPT-5.6 Sol Ultrafast with OpenAI.
493
SHARES
1.4k
VIEWS
Share on FacebookShare on Twitter

Cerebras published a bar chart with a 14x on it, and the bar beside it says 2.5x. That is the pitch for Ultrafast, the OpenAI API service tier previewed on 13 August: GPT-5.6 Sol running on Cerebras wafer scale hardware at up to 750 output tokens a second, up to fourteen times the Standard tier. Now scroll two images further down the same post. Cerebras also published its GDP-Val run, and there the gain is 5.6x end to end, because the part of the clock that is not the model, your tool calls plus your own code, went from 14.2 seconds to 14.9. It got very slightly slower. Credit where it is due, they published that chart themselves. It is a limited preview, for selected customers, and nobody has published a price.

The short answer

Ultrafast is a third service tier on the OpenAI API, running GPT-5.6 Sol on Cerebras wafer scale hardware. The headline is up to 14x the Standard tier. The measured end to end number on the vendor’s own benchmark is 5.6x, and the gap between those two numbers is the useful part. Limited preview, selected customers, nothing you can budget against yet.

750/speak output tokens, Cerebras measured
5.6xend to end on their own GDP-Val run
no priceand no quota, region or SLA
Answer card: OpenAI previewed the Ultrafast service tier on 13 August 2026, running GPT-5.6 Sol on Cerebras hardware at up to 750 output tokens per second and up to 14 times the Standard tier, in limited preview to selected customers with no published price, quota or service level.
A tier with a multiplier and no rate card.

What actually launched

Not a model. A lane.

Ultrafast is a service tier, sitting alongside Standard and Priority on the same OpenAI API, serving the same GPT-5.6 Sol you already call. Same weights, same interface, different silicon underneath: Cerebras wafer scale engines rather than GPUs, which is what the ten billion dollar partnership the two companies signed earlier this year was for. Cerebras keeps 44 GB of SRAM on the chip itself, so serving a token doesn’t mean hauling weights back and forth across a memory bus, and that single architectural fact is where the speed comes from.

Cerebras launch artwork titled Accelerating GPT-5.6 Sol Ultrafast with OpenAI, white text on a dark background crossed by orange and blue light trails. Image: Cerebras

Sol is the right model to put there, incidentally. It’s the slow one. The member of the GPT-5.6 family OpenAI points at legal briefs and engineering reports, the one that thinks for minutes, and if you’re going to spend money making inference fast then that’s where the seconds are.

The number under the number

Here’s the chart Cerebras published, and it’s more interesting than the one everybody quoted.

Diagram breaking down the Cerebras GDP-Val wall clock chart: Sol Ultrafast at 68.1 seconds of model request plus 14.9 seconds of other time for 83.0 seconds total, against Sol Standard at 7.5 minutes of model request plus 14.2 seconds for 7.7 minutes total, giving 6.6 times on the model request and 5.6 times end to end.
The grey tail is the same width in both bars. That's the ceiling.

On six quality matched GDP-Val tasks, Standard spends 7.5 minutes on the model request and 14.2 seconds on everything else. Ultrafast spends 68.1 seconds on the model request and 14.9 seconds on everything else. So the model request got 6.6 times faster, the whole task got 5.6 times faster, and the non inference part got marginally worse.

That gap isn’t a Cerebras problem, it’s just Amdahl’s law showing up on a launch page. Fourteen times is the number you’d see if a request were nothing but token generation. Real agent work is loops, retrieval and code you wrote, and none of that runs on a wafer.

Cerebras chart comparing inference and non inference wall clock across six quality matched GDP-Val tasks, showing GPT-5.6 Sol Ultrafast at 83.0 seconds total against GPT-5.6 Sol at 7.7 minutes total, labelled 5.6 times faster than Sol Standard. Image: Cerebras

The other measurement worth keeping is Humanity’s Last Exam, 2,500 questions, run in 11 hours 11 minutes against 78 hours 27 minutes for the comparison model. Roughly seven times. That one is a batch job, which is exactly the shape where a token generation win survives intact, and it’s a genuinely useful data point for anyone running large evaluation sweeps overnight.

All of these are Cerebras measuring hardware Cerebras sells, dated July 2026, with no independent re-run. We’d hold them loosely. The internal consistency is good though, and a vendor publishing the chart that undercuts its own headline number is not the usual behaviour.

Three rungs, two rate cards

Comparison chart of the three GPT-5.6 Sol service tiers by speed relative to Standard: Ultrafast at up to 14 times with no published price, Priority at 2.5 times with a published price, and Standard at 1 times as the priced baseline.
The middle column is where the news is.

Standard is 1x and priced. Priority is 2.5x and priced, at roughly double Standard by most accounts of the launch. Ultrafast is up to 14x and priced at whatever you negotiate, because OpenAI published no rate, no quota, no region list and no latency term with the preview.

Which makes the only question that matters unanswerable from outside. Nobody is buying tokens here, they’re buying seconds off a wall clock, and the metric is cost per second removed. If Priority buys 2.5x for 2x the money, Ultrafast has a bar to clear, and the shape of that bar is entirely unknown until a rate card exists. I might be wrong, but a tier that launches to Jane Street with no public price usually means the price is high enough to want a conversation first.

Named early users are Jane Street, Podium, Basis and Rogo. Trading, sales tooling, accounting and finance research. Latency sensitive money, in other words, which fits.

Who should care today

If you run interactive work where a human is waiting, this is the first tier that changes what the product can be, and worth joining the queue for. Voice, live research, an agent someone is watching.

If you run overnight batch, the HLE result says the win is real for you too and it may be the easiest sell, since you can measure it directly against your current bill.

And if your agent spends half its wall clock in tool calls, do the arithmetic on your own traces before you get excited. Take your model request time, divide by six, add back everything else unchanged. That’s your realistic ceiling, and for a lot of pipelines it lands nearer 2x than 14x.

Nothing changes in your code either way, which is the quiet good news. Same API, same model ID, a tier flag. That’s a very different migration from the GPT-5.6 Sol retune that split ChatGPT from the API earlier this month.

Worth noting where this sits in the wider picture. OpenAI is now buying inference speed from Cerebras wafers while also building its own inference accelerator with Broadcom. Both bets are on the same conviction: the model is fast enough, the serving is not.

Sources

Cerebras, Accelerating GPT-5.6 Sol Ultrafast with OpenAI, 13 August 2026, for the 750 tokens per second figure, the GDP-Val wall clock chart, the Humanity’s Last Exam timings and the tier multipliers. OpenAI, Previewing Ultrafast mode, 13 August 2026. Unite.AI, Cerebras runs OpenAI’s GPT-5.6 Sol at 750 tokens per second, for the named early customers and the on chip SRAM detail. The Decoder, GPT-5.6 Sol goes 14x faster as OpenAI launches Ultrafast mode, for the Priority tier pricing comparison and the partnership figure. Both charts reproduced above are Cerebras’ own, from the announcement post.

Frequently asked questions

What is OpenAI Ultrafast?

A new service tier on the OpenAI API, previewed on 13 August 2026, that runs GPT-5.6 Sol on Cerebras wafer scale hardware instead of GPUs. OpenAI puts it at up to 14 times the speed of the Standard tier, and Cerebras measured up to 750 output tokens per second. It is a limited preview open to selected customers, expanding as capacity allows.

How much does Ultrafast cost?

No rate has been published. OpenAI's preview announcement carries speed figures and named customers but no price, no quota, no region list and no latency commitment. The neighbouring Priority tier is the one you can already cost out, and coverage of the launch put it at roughly double Standard for 2.5 times the speed.

Is Ultrafast really 14 times faster?

For token generation against the Standard tier, on OpenAI's own chart, up to. On a real workload it is less. Cerebras' GDP-Val figures give 6.6x on model request time and 5.6x once the rest of the task is counted, because non inference time went from 14.2 to 14.9 seconds. The more tool calling a workload does, the closer it sits to that floor.

Which models run on Ultrafast?

GPT-5.6 Sol at launch, the reasoning heavy member of the GPT-5.6 family that OpenAI positions for legal, financial and engineering work. Terra and Luna are not part of the preview. Cerebras describes the tier as serving Sol specifically, and OpenAI has not said which models follow.

Who can use Ultrafast today?

Selected customers only, through the OpenAI API. Jane Street, Podium, Basis and Rogo have been named as early users. There is a sign up form for everyone else, and OpenAI says access widens as Cerebras capacity grows. There is no published date for general availability.

Tags: aiapiinferencellmnewsopenai
Share197Tweet123
stephane

stephane

  • Trending
  • Comments
  • Latest
Answer card: Proton Lumo 2.0 is private by policy, not by locality. Saved history is locked so even Proton cannot read it, but the prompt is decrypted on a Proton EU server to answer it, then forgotten.

Proton Lumo 2.0 review: how private is it, really?

3 September 2026
The Agentic Coding section of the official Hy4 preview benchmark appendix published by Tencent, a table comparing Hy3 and Hy4 preview against DeepSeek V4 Pro 0813, Qwen 3.8 Max, GLM 5.3, Kimi K3, GPT 5.6 Sol and Claude Opus 5 across SWE-bench Multilingual, SWE-bench Pro, DeepSWE, three SWE Atlas tasks, SWE-Marathon, Terminal-Bench 2.1, NL2Repo-Bench, CyberGym, ProgramBench, PostTrainBench and Harbor-Index.

Tencent’s 770B Hy4 tops one benchmark row in 46

3 September 2026
Answer card: Qwen 3.7 Max is API-only and cannot run locally yet; the open Qwen models (Qwen 3.6 27B, qwen3:8b to 32b) run offline via Ollama.

Qwen 3.7 local: what you can actually run offline

22 June 2026
Answer card: JWTs are not encrypted, anyone can read them; the signature proves who issued the token, not who may read it.

Are JWTs encrypted? No, and the difference will bite you

0
Answer card: a random 8 character password falls in under 2 hours offline, while 16 random characters hold for 1.4 trillion years at the same speed.

How long does it take to crack a password in 2026?

0
Answer card: three DNS records decide if your mail lands or bounces; SPF lists allowed senders, DKIM signs messages, DMARC sets the failure policy.

SPF, DKIM and DMARC explained: the records your email needs

0
The official xAI announcement card for Grok 4.7, white type on a dark grey and navy gradient.

Grok 4.7 keeps $2 and $6, and its gains over 4.6 are xhigh versus high

22 September 2026
Answer card stating that Qwen-Image-2.1, released on 20 September 2026, ships open weights with a 7 billion parameter diffusion transformer, a Qwen3-VL 8B text encoder and an RGBA VAE totalling about 33 gigabytes in BF16, under the Qwen Research License that limits use to research or evaluation and requires a separate commercial licence, unlike the Apache 2.0 licence of Qwen-Image 1.0.

Qwen-Image-2.1 brings the weights back, but not the Apache licence

21 September 2026
Answer card stating that Ternary Bonsai 2 27B, released by PrismML on 17 September 2026 under Apache 2.0, packs Qwen3.8 27B into 5.95 gigabytes at 1.72 bits per weight, keeps 98.2 percent of the 14-benchmark average, about 75 percent on SWE-bench Verified and Terminal-Bench 2.1, and needs PrismML's llama.cpp fork to run.

Does Bonsai 2 27B really keep 98% of Qwen3.8 in 5.95 GB?

20 September 2026
  • About
  • Contact
  • Privacy
  • Legal

Copyright © 2026 Stephane Cardon.

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • Home
  • Articles
    • Security
    • Network
    • Dev
    • Sysadmin
    • SEO
    • Email & DNS
  • Tools
    • Network tools: free, fast, no signup
    • Security tools: free, fast, no signup
    • Developer tools: free, fast, no signup
    • Sysadmin tools: free, fast, no signup
    • SEO tools: free, fast, no signup
    • Email & DNS tools: free, fast, no signup
  • Download
  • About

Copyright © 2026 Stephane Cardon.