DevNews

AMD Helios vs Vera Rubin NVL72: the MI455X numbers

On this page
  1. What AMD actually shipped
  2. The one claim that needs no benchmark
  3. The 30% is a model, not a measurement
  4. In production today, and pre-production in the footnote
  5. What to do with this

Lisa Su held a chip up on stage in San Francisco on July 23 and called Helios the best rack in the world today. The slide behind her read IN PRODUCTION TODAY. So here's the thing itself: a double-wide rack with 72 AMD Instinct MI455X GPUs and 18 sixth-gen EPYC Venice CPUs, 31 terabytes of HBM4, and a pitch of up to 15% more compute and 30% more tokens per dollar than an NVIDIA Vera Rubin NVL72. The hardware is real, and it's genuinely the first credible rack-scale answer AMD has built. Then we went and read AMD's own footnotes, which is where this gets interesting. The compute claim is peak theoretical math across two different FP4 formats. The tokens-per-dollar number leans on AMD's projected hourly pricing for its competitor's GPUs.

The short answer

AMD’s Helios rack is real hardware and a real answer to NVIDIA at rack scale. The memory story holds up: 432 GB of HBM4 per GPU, 31 TB per rack, roughly 50% more than a Vera Rubin NVL72, straight off both datasheets. The performance story is softer. AMD’s own footnotes call the compute win peak theoretical across mismatched FP4 formats, and the tokens-per-dollar headline uses projected pricing for a competitor’s GPUs. Here’s how to read each number.

July 23launched
72x MI455X31 TB HBM4
30%tokens/$, AMD estimate
Answer card: AMD Helios launched July 23 with 72 Instinct MI455X GPUs and 18 EPYC Venice CPUs, 31 TB of HBM4, 2.9 exaflops of dense FP4 and 260 TB/s of scale-up bandwidth, with AMD claiming up to 15 percent more compute and 30 percent more tokens per dollar than a Vera Rubin NVL72 based on Performance Labs calculations against NVIDIA published specifications.
The one-card version. Real rack, modeled comparison. PNG

What AMD actually shipped

Not a chip. A rack.

That’s the shift worth registering. AMD spent years selling Instinct GPUs that you or your OEM bolted into something, while NVIDIA sold the whole rack as one co-designed unit. Helios is AMD finally doing the same: 72 Instinct MI455X GPUs in a single scale-up domain, 18 sixth-gen EPYC “Venice” 9006 Series CPUs on the host layer, Pensando networking, all of it stitched with a UALoE fabric and running ROCm. AMD quotes 2.9 exaflops of dense FP4, 1.4 exaflops of FP8, 31 TB of HBM4 and 260 TB/s of scale-up bandwidth across the rack.

AMD launch slide showing a row of AMD Helios racks alongside the EPYC, Instinct, Pensando and ROCm logos, with the caption In Production Today. Image: AMD

The individual GPU is the interesting part. MI455X carries 432 GB of HBM4 at 23.3 TB/s, which is a lot of memory per accelerator by any 2026 standard. AMD also went with merchant silicon and open standards where NVIDIA went proprietary: UALink over Ethernet inside the rack, Ultra Ethernet Consortium between racks, Broadcom switch ASICs doing the scale-up. Whether that openness translates into better prices for buyers is the kind of thing we’ll only know in two years, but it’s a real strategic difference and not just a slide.

Customer names are unusually heavy for a launch: OpenAI, Anthropic, Meta, Microsoft, Oracle, plus AT&T and Cisco. Anthropic has committed to up to 2 gigawatts of MI455X in Helios racks. OpenAI expects to bring Helios online beginning in Q4 2026. Those are commitments, not running clusters, the same way the Vera Rubin deployment at Bristol Myers Squibb was a purchase order dressed as a data center.

The one claim that needs no benchmark

Start with the number that holds, because there is one.

Bar chart comparing HBM4 memory capacity per 72-GPU rack: AMD Helios at 31 TB versus NVIDIA Vera Rubin NVL72 at about 20.7 TB, roughly 50 percent more on the AMD side.
Memory capacity is stamped on both datasheets. This one is arithmetic, not a model. PNG

MI455X has 432 GB of HBM4 per GPU where the Rubin part has 288 GB. Multiply by 72 and you get 31 TB against roughly 20.7 TB. AMD’s footnote for this claim, MI400-007, says plainly that it compares published memory capacity and bandwidth specifications on both sides. No benchmark, no modeling, no interpretation. Capacity is capacity, and if your bottleneck is fitting a large model plus its KV cache inside one scale-up domain without sharding across racks, 50% more of it is worth exactly what it sounds like.

Same story for scale-out bandwidth (footnote MI400-019), where AMD quotes 43 TB/s per rack and again compares published specs. Honestly, this is where I’d focus if I were evaluating Helios. Memory capacity and fabric width are the two things you can verify before signing anything.

Bandwidth per GPU, worth noting, is nearly a tie: 23.3 TB/s on MI455X against 22 TB/s on the Rubin side. AMD doesn’t lead with that one.

The 30% is a model, not a measurement

Now the headline. “Up to 30% more tokens per dollar than the leading competitive solution” is the line AMD put in its press release, and it’s the line that made the coverage.

Footnote MI400-025 spells out what’s behind it. The figure comes from “AMD Performance Labs estimates as of July 2026”, calculated on the Kimi K2 Thinking workload at 32K input and 8K output, reflecting “estimated aggregate throughput” across low, medium and high interactivity, and using “projected hourly pricing for the system GPUs”.

Read that last clause twice. Tokens per dollar is a ratio, and AMD is projecting both halves. It doesn’t set the price of a Helios rack (Helios is a reference design, so Bull, HPE, Lenovo and Supermicro decide what you pay) and it obviously doesn’t set NVIDIA’s. An unconfirmed estimate attributed to Futurum research and circulating since July 21 puts Helios at 5 to 5.5 million dollars a rack against 3.5 to 4 million for Vera Rubin. If anything close to that holds, a 40% price premium eats a 30% efficiency edge and then some. I want to be careful here, because that estimate isn’t confirmed by anyone and rack pricing at gigawatt scale bears no resemblance to list. But it’s a live question that AMD’s own number can’t settle.

Checklist separating AMD Helios claims into datasheet arithmetic and modeled projections, noting that memory and scale-out comparisons use published specs while the compute win, throughput win and tokens-per-dollar figure come from AMD Performance Labs calculations, projections and pre-production hardware.
Sorted by what AMD's own footnotes say each number is. PNG

The throughput claim has the same shape. AMD says Helios delivers up to 15% higher tokens per GPU at low interactivity, 12% at medium and 10% at high, on Kimi K2 Thinking. Footnote MI400-023 says that was calculated “compared with published specifications for the NVIDIA Vera Rubin NVL72 rack”, and the blog body describes both sides as modeled. Nobody ran two racks side by side. AMD couldn’t have, really, and neither could anyone else yet.

Then there’s the compute headline, the 15% more AI compute. Footnote MI400-005 says it’s peak theoretical precision performance, and it compares AMD’s MXFP4 against NVIDIA’s dense NVFP4. Those are different four-bit formats with different scaling schemes. Peak flops in one doesn’t convert to peak flops in the other, which is why StorageReview’s teardown notes that MI455X sustains around half its 40.26 PFLOPS peak MXFP4 rating on real work anyway. Peak numbers were always soft. Peak numbers across mismatched formats are softer.

Also worth a note: the workload AMD picked is Kimi K2 Thinking, not the K3 that Moonshot shipped on July 16. Benchmarking against last generation’s open model is normal (K3 landed a week before the event) and I don’t think it’s a dodge. It just means the numbers describe a model most people have already moved past.

In production today, and pre-production in the footnote

The slide said IN PRODUCTION TODAY. AMD’s press release said “now in production to be deployed by leading AI companies at gigawatt scale.”

At the bottom of the same blog post: “Performance measurements were obtained on pre-production or reference hardware under specific workload and configuration conditions” and “Figures are projected, subject to change, and do not represent a commitment regarding final specifications.”

Both things can be true. Racks can be coming off a line while the numbers on the slides came from pre-production silicon. But the gap between the marketing tense and the legal tense is the single most useful signal in this launch, and it puts Helios in the same place as most 2026 AI hardware: announced, committed to, going online next year. OpenAI’s Q4 2026 date is the earliest concrete one on the board.

What to do with this

If you’re specifying a cluster, three of AMD’s numbers are checkable today and you should check them yourself: HBM4 capacity per rack, scale-out bandwidth, and power draw. That last one nobody advertises. StorageReview puts a reference Helios at 225 to 245 kW depending on workload, which is a genuinely large number in a world where New York started gating data centers above 50 MW and where power, not silicon, decides whether your build happens.

Everything else, wait for it. Not because AMD is being dishonest (the footnotes are unusually thorough, and they’re how we wrote most of this piece) but because a vendor modeling its competitor’s rack from a spec sheet is doing the best it can with what it has, which is not the same as a measurement. The first independent Helios-versus-Rubin run will tell you more than every number above.

For now the honest summary is short. AMD built a real rack with more memory in it. Whether it serves tokens cheaper is unproven, and the price nobody will publish decides it.

Sources: AMD blog, “AMD Launches Helios”, AMD press release, July 23 2026, StorageReview, TechCrunch and Duckit Tech on the Futurum price estimate, July 2026. The rack configuration, the 2.9 and 1.4 exaflops figures, 31 TB of HBM4, 260 TB/s scale-up and 43 TB/s scale-out, the 15% compute and 30% tokens-per-dollar claims and footnotes MI400-005, MI400-007, MI400-019, MI400-023 and MI400-025 are quoted from AMD’s own launch blog and press release. Per-GPU HBM4 capacity and bandwidth for both vendors, the 20.7 TB Vera Rubin rack total, the sustained-MXFP4 observation and the 225 to 245 kW reference draw are from StorageReview’s analysis. The 5 to 5.5 million dollar Helios and 3.5 to 4 million dollar Vera Rubin rack prices are an analyst estimate attributed to Futurum, unconfirmed by either vendor, and should be treated as such. The image is AMD’s own launch graphic.

Frequently asked questions

What is AMD Helios?

A rack-scale AI system AMD launched on July 23, 2026 at its Advancing AI conference. One Helios rack connects 72 Instinct MI455X GPUs and 18 sixth-gen EPYC "Venice" 9006 Series CPUs over AMD Pensando networking and a UALoE fabric, with ROCm as the software layer. AMD quotes 2.9 exaflops of dense FP4, 1.4 exaflops of FP8, 31 TB of HBM4 and 260 TB/s of scale-up bandwidth. Systems ship through Bull, HPE, Lenovo and Supermicro.

Is Helios really faster than NVIDIA Vera Rubin NVL72?

On paper, on some axes, according to AMD. The company claims up to 15% more AI compute, 50% more HBM capacity and 50% more scale-out bandwidth. The memory and scale-out numbers come from published datasheets on both racks, so you can redo that arithmetic. The compute figure is peak theoretical performance comparing AMD MXFP4 against NVIDIA NVFP4, which are different formats. Nobody outside AMD has benchmarked the two racks against each other.

What does the 30% more tokens per dollar claim actually rest on?

AMD footnote MI400-025. It says the figure comes from AMD Performance Labs estimates using the Kimi K2 Thinking workload at 32K input and 8K output, comparing aggregate throughput across three interactivity points, and using projected hourly pricing for the system GPUs. AMD does not set its own rack price (Helios is a reference design that OEMs build) and it certainly does not set NVIDIA's. Treat the dollar half of tokens-per-dollar as an assumption.

How much does a Helios rack cost?

AMD has not published a price. An estimate attributed to Futurum research and circulated on July 21 put Helios at roughly 5 to 5.5 million dollars per rack against 3.5 to 4 million for a Vera Rubin NVL72, which would be about 40% more. That is unconfirmed by AMD and worth holding loosely, because final pricing depends on the OEM, memory costs, networking configuration and contract scale.

Can I buy one today?

Not off a web page. AMD says Helios is in production for gigawatt-scale deployment with OpenAI, Anthropic, Meta, Microsoft, Oracle and others, and OpenAI expects to bring Helios online starting in the fourth quarter of 2026. AMD's own disclaimer notes that its performance measurements were taken on pre-production or reference hardware. If you are not buying by the gigawatt, this is a 2027 conversation with an OEM.