• Latest
  • Trending
  • All
Close-up photograph of the Jalapeño package mounted on a teal green test board, a large square heat spreader frame around an exposed silicon die, surrounded by rows of gold capacitors and connectors, with a Broadcom logo visible at the bottom edge of the board.

Jalapeño measures 1.9x per watt against a GB200

3 September 2026
The official xAI announcement card for Grok 4.7, white type on a dark grey and navy gradient.

Grok 4.7 keeps $2 and $6, and its gains over 4.6 are xhigh versus high

22 September 2026
Answer card stating that Qwen-Image-2.1, released on 20 September 2026, ships open weights with a 7 billion parameter diffusion transformer, a Qwen3-VL 8B text encoder and an RGBA VAE totalling about 33 gigabytes in BF16, under the Qwen Research License that limits use to research or evaluation and requires a separate commercial licence, unlike the Apache 2.0 licence of Qwen-Image 1.0.

Qwen-Image-2.1 brings the weights back, but not the Apache licence

21 September 2026
Answer card stating that Ternary Bonsai 2 27B, released by PrismML on 17 September 2026 under Apache 2.0, packs Qwen3.8 27B into 5.95 gigabytes at 1.72 bits per weight, keeps 98.2 percent of the 14-benchmark average, about 75 percent on SWE-bench Verified and Terminal-Bench 2.1, and needs PrismML's llama.cpp fork to run.

Does Bonsai 2 27B really keep 98% of Qwen3.8 in 5.95 GB?

20 September 2026
Answer card stating that Jev 1.13 from TypeSafe AI is a decision model in early access since 15 September 2026 that returns typed probabilities instead of text, priced at 42 dollars per billion input tokens with output tokens free, answering in 70 to 500 milliseconds, with a 64K token request budget, text input only, and a documented list of things it does badly, including counting and dates.

Jev 1.13 bills $42 a billion tokens, and it can’t count

19 September 2026
Answer card stating that Qwen3.8-Omni-Flash launched on 17 September 2026 as an API only model on Alibaba Cloud Model Studio, taking text, images, audio and video in a 1M token context and returning text only, priced at 0.15 dollars per million input tokens for every modality and 0.47 dollars per million output tokens in the international regions, with no open weights published and the Qwen-Live Harness GitHub repository returning 404.

Qwen3.8-Omni-Flash bills audio at $0.15 and ships no weights

18 September 2026
Answer card stating that on 15 September 2026 AWS said it is unable to restore access to resources and data hosted exclusively in the Middle East Bahrain region me-south-1 and in the mec1-az2 zone of the UAE region, because the damage spanned multiple Availability Zones and exceeded what multi-AZ services are designed to withstand.

AWS can’t restore me-south-1, six months after the drone strikes

17 September 2026
Answer card stating that Google released Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking on 15 September 2026 at 3 dollars per million audio input tokens and 12 dollars out, that the thinking model requires asynchronous tools, and that Artificial Analysis scores it 82.6 on its Speech to Speech Quality Index.

Gemini 3.8 Live Extended Thinking rejects any tool that blocks

16 September 2026
Answer card summarising the Atria Dawn Preview release: 744B GLM-5.2 base, MIT licence, 1.5 TB BF16 and 756 GB FP8 checkpoints, 256K context, top on five of sixteen benchmark rows and trailing on SWE-bench Pro.

Atria Dawn Preview is 744B under MIT, and the BF16 weighs 1.5 TB

15 September 2026
Answer card stating that OpenAI released the Agents API in public beta on 10 September 2026 with no separate fee, billed through model tokens, tool calls and hosted sandbox time, with a choice of OpenAI hosted, self hosted or partner sandboxes, US only data residency and no Zero Data Retention support.

OpenAI’s Agents API has no fee, no ZDR and a one hour sandbox clock

14 September 2026
Answer card: Sakana Fugu Max at $2 and $6 per million tokens, Fugu Ultra v2 unchanged at $5 and $30, and Sakana saying Ultra v2 scores without Fable 5 or GPT-6 Astra in its pool.

Fugu Max costs $2 and $6 while Fugu Ultra v2 runs without Fable 5

13 September 2026
Answer card stating that DeepSeek released DeepSeek-V4.1-Flash on 10 September 2026 as a 552 billion parameter mixture of experts model with a new causal encoder decoder architecture that activates 8 billion parameters on input and 16 billion on output, with native vision, a one million token context and MIT licensed weights, that the API model name is now deepseek-flash at 0.15 dollars per million input tokens and 0.60 dollars per million output tokens off peak, and that DeepSeek announced V4 Pro would be routed to V4.1-Flash from 14 September and reversed that on 11 September.

DeepSeek V4.1-Flash arrived, and the V4 Pro retirement lasted a day

12 September 2026
Answer card stating that Cognition released SWE-2 on 10 September 2026, a coding model post-trained from Kimi K3, scoring 50.0 percent on FrontierCode 1.1 Main against 50.9 percent for Claude Fable 5.1 and 27.3 percent on Terminal-Bench 4 against 55.8 percent, available only inside Devin.

SWE-2 trails Fable 5.1 by one point, and by 28 on Terminal-Bench 4

11 September 2026
  • About
  • Contact
  • Privacy
  • Legal
Wednesday, September 23, 2026
  • Login
Packet Nebula
  • Home
  • Articles
    • Security
    • Network
    • Dev
    • Sysadmin
    • SEO
    • Email & DNS
  • Tools
    • Network tools: free, fast, no signup
    • Security tools: free, fast, no signup
    • Developer tools: free, fast, no signup
    • Sysadmin tools: free, fast, no signup
    • SEO tools: free, fast, no signup
    • Email & DNS tools: free, fast, no signup
  • Download
  • About
No Result
View All Result
Packet Nebula
No Result
View All Result
Home Dev

Jalapeño measures 1.9x per watt against a GB200

by stephane
3 September 2026
in Dev
0
Close-up photograph of the Jalapeño package mounted on a teal green test board, a large square heat spreader frame around an exposed silicon die, surrounded by rows of gold capacitors and connectors, with a Broadcom logo visible at the bottom edge of the board.
497
SHARES
1.4k
VIEWS
Share on FacebookShare on Twitter

A photograph of a chip sitting on a teal test board, and finally some numbers under it. OpenAI published Jalapeño's first measured results on 25 August, and the short version is that its custom inference chip did between 1.5 and 1.9 times more work per watt than the Nvidia systems it ran against, with end to end latency 1.7 to 3.6 times lower. Those runs look real. They sit on InferenceX, a public benchmark from SemiAnalysis, across three open weight models anyone can download. What deserves a second look is the other half of the matchup. The comparison hardware is a GB200 in one chart and a GB300 in the other two, so this is Blackwell, and SemiAnalysis, which owns the benchmark, says out loud that Jalapeño should have been put up against Vera Rubin instead.

The short answer

OpenAI posted the first measured results for Jalapeño on 25 August 2026, seven weeks after showing the chip itself. Runs on InferenceX, a public benchmark from SemiAnalysis, at a nominal 8k in and 1k out, across GPT-OSS 120B, DeepSeek R1 670B and Kimi K2.5 1T. Peak throughput per kilowatt comes out 1.5x to 1.9x ahead, end to end latency 1.7x to 3.6x lower. The appendix names the comparison hardware as a GB200 and a GB300, both Blackwell. SemiAnalysis says the right opponent would have been Vera Rubin.

1.9xwork per watt, GPT-OSS 120B vs GB200
3.6xlower latency, DeepSeek R1 vs GB300
0units for sale, rent or independent testing
Answer card stating that OpenAI published the first measured results for its Jalapeno inference chip on 25 August 2026, run on the InferenceX benchmark from SemiAnalysis, showing 1.5 to 1.9 times more work per watt at peak throughput and 1.7 to 3.6 times lower end to end latency than the Nvidia GB200 and GB300 systems used for comparison.
Seven weeks ago this chip had no published numbers. Now it has an appendix.

The numbers, and the footnote under them

When we wrote about Jalapeño in July, the honest summary was that OpenAI had a package, a partner in Broadcom and a nine month tape-out story, with nothing measured. That gap closed on Tuesday at Hot Chips.

The figures worth keeping are in the appendix rather than the headline. On GPT-OSS 120B, Jalapeño hit 85,448 mixed tokens per second per kilowatt against 44,960 for the GB200. On DeepSeek R1 it was 19,641 against 11,781 for a GB300, and on Kimi K2.5 18,195 against 11,862. Latency told a bigger story than throughput: 1.65 seconds against 5.99 on DeepSeek R1, 1.56 against 5.31 on Kimi.

Bar chart of peak mixed tokens per second per kilowatt on the InferenceX benchmark comparing Jalapeno with the Nvidia GB200 on GPT-OSS 120B at 85448 against 44960, and with the Nvidia GB300 on DeepSeek R1 670B at 19641 against 11781 and on Kimi K2.5 1T at 18195 against 11862, all normalised on published package TDP.
Two of these bars are a GB300. One is a GB200. The chart headline never says so.

Per kilowatt, not per chip. OpenAI is direct about preferring that denominator, and we think it’s the right one for anybody who pays a power bill. It also normalises each system on its published package TDP rather than on what the thing actually drew. Jalapeño is rated at 700 W and OpenAI says it stayed at or below 550 W across these workloads, so on its own chosen metric the chip is being handicapped by roughly a quarter. That’s a slightly unusual thing for a vendor to volunteer.

The opponent is a generation behind

Here’s the part that changes how you read all of it.

SemiAnalysis publishes InferenceX and verified these runs in person in its lab. It also wrote, in the same week, that the Blackwell matchup is incomplete and unfair, and that Jalapeño really ought to be measured against Vera Rubin, which uses the same HBM4 memory generation. Its blunt phrasing was that a custom chip like this is expected to beat Blackwell.

That is not a debunking. It’s a scale correction, and it comes from the people who own the yardstick, which is about as good as sourcing gets. Vera Rubin systems are already going into AI factories, so by the time Jalapeño moves any real volume it will be sharing racks with the generation it was never benchmarked against. Richard Ho, who runs hardware at OpenAI, said as much himself: by full deployment the competition may have moved on.

One more methodology note that nobody will read. Every run here is single turn, 8k in and 1k out. SemiAnalysis says its preferred suite for comparing chips is AgentX, because multi turn and long context work is what actually stresses routers, prefix caches and offload paths. No AgentX results were published. For a chip whose entire pitch is agentic serving, that’s the missing table.

About those 104x numbers

The appendix carries some figures that look like typos. At the previous best time between tokens, Jalapeño delivers 104.3 times more throughput per kilowatt than the GB300 on DeepSeek R1. On GPT-OSS against the GB200 it’s 53.7 times.

Both are real arithmetic and both are close to meaningless out of context. They measure throughput at a fixed latency target the Nvidia system can barely reach, so the comparison is taken at the exact point where the older architecture falls off a cliff. Push any accelerator down to its minimum time between tokens and efficiency collapses. Quote the 1.5x to 1.9x instead. That’s the number that survives contact with a rack.

Checklist separating what the 25 August 2026 Jalapeno results establish, namely measured runs on a public third party benchmark across three open weight models with the comparison hardware named and the power normalisation disclosed, from what they leave open, including the absence of a Vera Rubin comparison, no AgentX multi turn results and nothing available to rent.
Three things settled. Two that a follow-up post will have to answer.

What the silicon looks like

SemiAnalysis published specs alongside its analysis, and they fill in what the July announcement left blank. HBM4 at 15.4 TB/s of bandwidth, running 10 Gbps pin speeds. Compute die on TSMC N3P, the I/O chiplet on N3E. The B0 stepping does 13.4 PFLOPs of MXFP4 inside that 700 W envelope.

The packaging is where the design argument lives. 128 Jalapeño ASICs per rack, arranged as 16 trays of 8, scaling out to 2,048 across 16 racks. OpenAI’s own description of the architecture keeps circling back to one idea: keep the KV cache local, keep the whole request inside a single connected network domain, and stop paying for data movement between prefill and decode. Whether that holds up on a 200k token agent loop is precisely what the missing AgentX run would have told us.

What actually reaches you

Nothing, directly. There’s no card, no instance type and no price. OpenAI starts deploying inside its own infrastructure by the end of this year in very small volumes, with 2027 as the real rollout, and it went out of its way to say it will keep buying Nvidia widely for training and inference both.

So the effect on your work is second order and slow. If Jalapeño does what the appendix says at scale, OpenAI’s cost per served token drops on its own models, and some of that eventually shows up as API pricing or as latency you can feel in an agent loop. I’d hold off on modelling any of it until Gen 2, honestly. First generation custom silicon has a habit of shipping late and quiet.

The part we’ll actually be watching is smaller and more interesting. OpenAI says AI generated implementations of selected GPT-OSS attention and mixture of experts blocks ran 1.5 to 1.8 times faster than the ones its human experts wrote, and that Codex with GPT-Astra brought three unplanned open weight models up to high performance in two months. Kernel work is the traditional reason custom accelerators die in the field, because nobody can afford to hand write support for every new model family. If that bottleneck is genuinely gone, it matters more than any bar chart on this page.

Close-up photograph of the Jalapeño package mounted on a teal green test board, a large square heat spreader frame around an exposed silicon die, surrounded by rows of gold capacitors and connectors, with a Broadcom logo visible at the bottom edge of the board.

Image: OpenAI, from its Jalapeño first results post, 25 August 2026.

Sources

All per model figures, the power normalisation method, the 700 W rating and the 550 W measured draw, the 2.1x to 4.1x interactive range and the AI generated kernel claims come from OpenAI’s own Jalapeño first results post and its appendix, published 25 August 2026. The Hot Chips context, the deployment volumes and the Richard Ho comments are reported by TechCrunch. The InferenceX methodology caveats, the AgentX point, the Vera Rubin argument and the hardware specifications are from SemiAnalysis, which publishes the benchmark and says it verified the runs in its lab.

Frequently asked questions

What did OpenAI actually measure with Jalapeño?

OpenAI ran Jalapeño on InferenceX, a public inference benchmark from SemiAnalysis, using GPT-OSS 120B, DeepSeek R1 670B and Kimi K2.5 1T at a nominal 8k input and 1k output. Across the three, it reported 1.5 to 1.9 times more work per watt at peak throughput and 1.7 to 3.6 times lower end to end latency than the comparison systems, rising to 2.1 to 4.1 times higher performance on highly interactive workloads.

What hardware was Jalapeño compared against?

OpenAI's post says only leading commercially available AI systems, but the appendix names them. GPT-OSS 120B was run against a GB200 at a 1,200 W package TDP, and both DeepSeek R1 and Kimi K2.5 against a GB300 at 1,400 W. Jalapeño is rated at 700 W. All three are Nvidia Blackwell generation parts, and Vera Rubin does not appear anywhere in the results.

Are the per watt figures fair?

They are transparent, which is not quite the same thing. OpenAI normalised every system on its published package TDP rather than on measured draw, and it says Jalapeño's sustained power stayed at or below 550 W on these workloads, so its own chip is scored on about 27 percent more power than it pulled. The bigger question is generational: SemiAnalysis, which publishes InferenceX, describes the Blackwell comparison as incomplete and unfair, because a chip taping out now should be measured against Vera Rubin.

Can I buy or rent a Jalapeño?

No. It is internal silicon for OpenAI's own data centres, with no card, no cloud instance and no published price. OpenAI plans to begin deploying it inside its own infrastructure by the end of 2026, in what its hardware lead described as very small volumes, with a more significant rollout during 2027. Generation 2 is in development and generation 3 is being designed.

Does this mean OpenAI stops buying Nvidia?

Not according to OpenAI. The same post states it will continue to widely deploy accelerators from Nvidia and other partners for both training and inference. Jalapeño targets serving cost on OpenAI's own workloads, so the realistic effect on anyone outside the company is cheaper or faster API responses at some point, not a change in what hardware they can rent.

Tags: aihardwareinferencenewsnvidiaopenai
Share199Tweet124
stephane

stephane

  • Trending
  • Comments
  • Latest
Answer card: Proton Lumo 2.0 is private by policy, not by locality. Saved history is locked so even Proton cannot read it, but the prompt is decrypted on a Proton EU server to answer it, then forgotten.

Proton Lumo 2.0 review: how private is it, really?

3 September 2026
The Agentic Coding section of the official Hy4 preview benchmark appendix published by Tencent, a table comparing Hy3 and Hy4 preview against DeepSeek V4 Pro 0813, Qwen 3.8 Max, GLM 5.3, Kimi K3, GPT 5.6 Sol and Claude Opus 5 across SWE-bench Multilingual, SWE-bench Pro, DeepSWE, three SWE Atlas tasks, SWE-Marathon, Terminal-Bench 2.1, NL2Repo-Bench, CyberGym, ProgramBench, PostTrainBench and Harbor-Index.

Tencent’s 770B Hy4 tops one benchmark row in 46

3 September 2026
Answer card: Qwen 3.7 Max is API-only and cannot run locally yet; the open Qwen models (Qwen 3.6 27B, qwen3:8b to 32b) run offline via Ollama.

Qwen 3.7 local: what you can actually run offline

22 June 2026
Answer card: JWTs are not encrypted, anyone can read them; the signature proves who issued the token, not who may read it.

Are JWTs encrypted? No, and the difference will bite you

0
Answer card: a random 8 character password falls in under 2 hours offline, while 16 random characters hold for 1.4 trillion years at the same speed.

How long does it take to crack a password in 2026?

0
Answer card: three DNS records decide if your mail lands or bounces; SPF lists allowed senders, DKIM signs messages, DMARC sets the failure policy.

SPF, DKIM and DMARC explained: the records your email needs

0
The official xAI announcement card for Grok 4.7, white type on a dark grey and navy gradient.

Grok 4.7 keeps $2 and $6, and its gains over 4.6 are xhigh versus high

22 September 2026
Answer card stating that Qwen-Image-2.1, released on 20 September 2026, ships open weights with a 7 billion parameter diffusion transformer, a Qwen3-VL 8B text encoder and an RGBA VAE totalling about 33 gigabytes in BF16, under the Qwen Research License that limits use to research or evaluation and requires a separate commercial licence, unlike the Apache 2.0 licence of Qwen-Image 1.0.

Qwen-Image-2.1 brings the weights back, but not the Apache licence

21 September 2026
Answer card stating that Ternary Bonsai 2 27B, released by PrismML on 17 September 2026 under Apache 2.0, packs Qwen3.8 27B into 5.95 gigabytes at 1.72 bits per weight, keeps 98.2 percent of the 14-benchmark average, about 75 percent on SWE-bench Verified and Terminal-Bench 2.1, and needs PrismML's llama.cpp fork to run.

Does Bonsai 2 27B really keep 98% of Qwen3.8 in 5.95 GB?

20 September 2026
  • About
  • Contact
  • Privacy
  • Legal

Copyright © 2026 Stephane Cardon.

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • Home
  • Articles
    • Security
    • Network
    • Dev
    • Sysadmin
    • SEO
    • Email & DNS
  • Tools
    • Network tools: free, fast, no signup
    • Security tools: free, fast, no signup
    • Developer tools: free, fast, no signup
    • Sysadmin tools: free, fast, no signup
    • SEO tools: free, fast, no signup
    • Email & DNS tools: free, fast, no signup
  • Download
  • About

Copyright © 2026 Stephane Cardon.