DevNews

AWS and Nvidia: 2 million GPUs, NVHBM inside Trainium4

On this page
  1. What went out on 26 August
  2. NVHBM, and what it does to the Trainium story
  3. The part that deserves a second read
  4. The memory bill explains the timing
  5. What you can actually book

Two million GPUs is the number in every headline, and honestly it's the least interesting thing in the announcement. AWS and Nvidia published a joint expansion on 26 August: Blackwell Ultra, Rubin and Rubin Ultra parts landing across AWS through 2027 and 2028, on top of the million-plus AWS already committed to at GTC. Fine. Buried further down is the line that made us stop reading. Annapurna Labs will build Trainium4 with Nvidia's NVHBM memory and NVLink Fusion. Amazon's own accelerator, the chip whose entire pitch was not paying the Nvidia tax, is going to ship with Nvidia's memory architecture inside it and share a rack with Nvidia GPUs. That's a different relationship than the one either company has been describing for the last three years.

The short answer

AWS and Nvidia announced a joint expansion on 26 August. The headline is two million extra GPUs across 2027 and 2028. The substance is further down: Annapurna Labs is putting Nvidia’s NVHBM memory and NVLink Fusion into next-generation Trainium, starting with Trainium4, so Amazon’s own accelerator and Nvidia GPUs land in a shared rack architecture. NVHBM relocates the memory controller into the HBM base die. Nvidia claims 30 percent more bandwidth per stack over HBM4e and 15 percent less HBM power. Nothing here ships this year except G7 instances.

2 millionmore Nvidia GPUs, 2027-2028
Trainium4first chip in line for NVHBM
30%bandwidth claim per stack, vendor figure
Answer card stating that AWS and Nvidia announced on 26 August 2026 that AWS will deploy two million additional Nvidia GPUs across 2027 and 2028 covering Blackwell Ultra, Rubin and Rubin Ultra, and that Annapurna Labs will build next-generation Trainium starting with Trainium4 using Nvidia NVHBM memory and NVLink Fusion so Trainium and Nvidia GPUs share a rack scale architecture.
The announcement in one card. The GPU count travelled further than the memory news, which is backwards. PNG

What went out on 26 August

The press release is long, so here’s the useful shape of it.

Two million additional Nvidia GPUs across AWS infrastructure in 2027 and 2028. Blackwell Ultra, Rubin and Rubin Ultra. AWS committed to more than a million at GTC 2026 and says demand exceeded what it planned for, which is the sort of line you’d expect either way but which does match a year of capacity being the binding constraint on everything.

Vera CPUs are coming to AWS. No date. Nvidia frames Vera as a host CPU for accelerated systems and a standalone option for agentic workloads, meaning the orchestration and tool-execution side that sits between model calls and eats more CPU than people budget for.

There’s a government piece: 100,000 GPUs on secure AWS infrastructure for federal workloads at Impact Level 6 and above. Nemotron models get first-party placement on Bedrock and SageMaker. Amazon Robotics is adopting Jetson, Omniverse and Isaac for simulation and synthetic data.

And then, without much ceremony, the Trainium line.

NVHBM, and what it does to the Trainium story

Nvidia published NVHBM the same day. It’s a memory architecture, and the idea is simpler than the acronym suggests.

On a normal accelerator, the memory controller lives on the compute die. HBM stacks sit next to it on the interposer, and a wide physical interface connects the two. That PHY costs die area, and die area is the most expensive real estate in the building. NVHBM moves the controller off the compute die and into the base die of the 3D HBM stack itself.

Diagram comparing a standard HBM4e layout, where the memory controller and a wide physical interface sit on the XPU compute die beside the memory stacks, with the NVHBM layout, where the memory controller moves into the base die of the 3D HBM stack, narrowing the interface and freeing compute die area.
One block moves. That's the whole architecture change, and it's why the numbers move in three directions at once. PNG

The claimed consequences: up to 30 percent more bandwidth per stack over HBM4e, up to 15 percent lower HBM power, and up to 25 percent more usable compute area because the interface shrinks. StorageReview, working from Nvidia’s material, puts the PHY support area reduction at 67 percent. Nvidia’s own scaling example is a one gigawatt site running 2,000 watt accelerators, where the power saved would fund roughly 15,000 more of them.

Treat every one of those numbers as a vendor claim. There is no shipping NVHBM part, no third-party measurement, and no date. I’d expect the real figures to land lower, because they always do, though the direction is clearly right: getting the controller out of the compute die is a well understood win and the reason nobody did it before is that it requires the memory vendors to build a custom base die for you.

Which is the actual interesting bit. Nvidia says NVHBM will be a standard implementation offered by multiple memory suppliers, not a one-vendor part. Nvidia isn’t making memory. It’s writing the spec and licensing it into other people’s accelerators.

Amazon is first. Nafea Bshara, VP at Annapurna Labs, is quoted calling NVHBM “a new architectural approach to advancing high-bandwidth memory performance”, which is press-release language, but his name being on it is the signal.

The part that deserves a second read

Trainium existed because AWS wanted a price floor that Nvidia didn’t set.

That was the pitch from the start, and customers understood it that way. Trainium4 will now use Nvidia’s memory architecture and Nvidia’s NVLink Fusion fabric, sitting in a rack architecture shared with Nvidia GPUs. AWS still designs the compute. Annapurna still owns the part it differentiates on.

But the two things that got expensive over the past eighteen months are memory and interconnect, and both of those now have Nvidia’s name on them inside Amazon’s chip.

Nvidia render of an AI accelerator package: a silver lid frame around a grid of dark compute dies, flanked by twelve gold high bandwidth memory stacks on either side, with interposer routing visible across the substrate.
Image: Nvidia (accelerator package with HBM stacks, from the NVHBM announcement)

I might be reading too much into it. A shared rack standard is genuinely good engineering, and a customer who can mix Trainium and Rubin in one scale-up domain has more options than one who can’t. Interoperability usually helps the buyer.

Still. When your escape hatch is built partly from the vendor you were escaping, the word independence is doing less work than it used to.

The memory bill explains the timing

None of this is happening in a vacuum, and the context is four days older than the announcement.

Bloomberg reported on 22 August that contract server builders had told Microsoft, Google and Oracle to expect price increases above 15 percent on Vera Rubin and Grace Blackwell systems from early 2027. TrendForce put memory at roughly 25 to 30 percent of an AI server’s bill of materials, up from 5 to 10 percent, which works out around two million dollars per rack. GPUs used to be more than 80 percent of the cost of these machines. They’re now about half that share.

So Nvidia is raising system prices because memory got expensive, and in the same week is licensing a memory architecture that claims to make memory cheaper per unit of bandwidth. Both things are true and they’re the same strategy. We’ve been tracking this crunch from the consumer end for months, from the Pixel 11 price increase onward. This is the data-centre end of the same shortage, and it’s where the money actually is.

Checklist separating what a team can act on today from what is 2027 or later in the AWS and Nvidia announcement: EC2 G7 instances with RTX PRO 4500 Blackwell are the only shipping item, the two million GPUs are 2027 and 2028 capacity, Trainium4 with NVHBM has no ship date, Vera CPU instances have no date, and the 100,000 GPU government build is a plan not an availability zone.
Sorted by whether you can put it in a 2026 budget. Most of it, you can't. PNG

What you can actually book

One thing.

EC2 G7 with the RTX PRO 4500 Blackwell Server Edition, which Nvidia rates at up to 4.6 times the AI inference throughput of G6 and 2.1 times the graphics throughput. That’s a real instance family with a real part in it. If you’re serving mid-sized models or doing rendering, it’s worth pricing against your current G6 spend.

Everything else is 2027 and 2028. The two million GPUs are capacity, not availability. Trainium4 has no ship date. Vera instances have no date. The government AI factory is a plan.

Two adjacent AWS numbers did come with the release and are worth noting if you run data pipelines: GPU-accelerated processing on Amazon EMR at 3.7 times faster with 30 percent better price-performance, and vector indexing on OpenSearch at 9 times faster for a quarter of the cost. Those are AWS’s figures on AWS’s benchmarks, so discount accordingly, but vector indexing cost is a real line item for anyone running retrieval at scale and a quarter is a big enough claim to test.

If you’re planning 2027 capacity, the useful takeaway isn’t the GPU count. It’s that memory pricing is now the variable that moves your quote, and that the Vera Rubin platform you’re budgeting for is going to cost more than the pre-announcement modelling said. Build the 15 percent in. If it lands lower, you’ll be pleasantly surprised, and nobody ever got fired for that.

Sources: the GPU numbers, Vera CPU plans, Trainium integration, government build, Nemotron availability, robotics adoption and the Garman and Huang quotes come from the Nvidia newsroom release of 26 August 2026 and Amazon’s own write-up. NVHBM architecture details, the bandwidth, power and die-area claims and the Bshara quote are from the Nvidia NVLink Fusion and NVHBM blog post, with the PHY area figure, the one gigawatt scaling example and the Trainium4 detail from StorageReview. The AI server price increase, the memory share of the bill of materials and the Bloomberg attribution are from TrendForce.

Frequently asked questions

What did AWS and Nvidia actually announce on 26 August 2026?

Two million additional Nvidia GPUs deployed across AWS infrastructure during 2027 and 2028, covering Blackwell Ultra, Rubin and Rubin Ultra parts. That sits on top of the more than one million GPUs AWS committed to at GTC 2026, a commitment AWS says demand ran through faster than expected. The release also covers Nvidia Vera CPUs coming to AWS, NVHBM and NVLink Fusion in next-generation Trainium, 100,000 GPUs for US government AI factories at Impact Level 6 and above, Nemotron models on Bedrock and SageMaker, and Amazon Robotics adopting Jetson, Omniverse and Isaac.

What is NVHBM and how is it different from HBM4e?

NVHBM moves the memory controller off the accelerator die and into the base die of the 3D HBM stack itself. Nvidia claims up to 30 percent more memory bandwidth per stack than HBM4e, up to 15 percent lower HBM power, and up to 25 percent more usable compute area on the accelerator die because the physical interface shrinks. Those are all vendor figures with no shipping silicon behind them yet. Nvidia says the standard will be offered by multiple memory suppliers rather than a single one.

Does NVHBM in Trainium4 mean AWS is giving up on its own silicon?

No, and reading it that way misses the point. Trainium4 is still an Annapurna Labs design and AWS still controls the compute architecture. What changes is the memory subsystem and the scale-up fabric, which now come from Nvidia. Amazon keeps the part it differentiates on and buys the part that got expensive. It is a narrower kind of independence than the original Trainium pitch implied, but it is not a retreat from custom silicon.

Can I use any of this today?

Almost none of it. The GPU numbers are 2027 and 2028 capacity, Trainium4 has no public ship date, and Vera CPU instances have no date either. The one thing with a real product attached is EC2 G7 with the RTX PRO 4500 Blackwell Server Edition, which Nvidia quotes at up to 4.6 times the AI inference throughput and 2.1 times the graphics throughput of G6. If your 2027 budget depends on capacity you can book, this announcement is not that.

How does this connect to the AI server price increases?

Directly. Bloomberg reported on 22 August that contract server builders warned Microsoft, Google and Oracle of price rises above 15 percent on Vera Rubin and Grace Blackwell systems from early 2027, driven by memory. TrendForce put memory at roughly 25 to 30 percent of an AI server bill of materials, up from 5 to 10 percent, around two million dollars per rack. NVHBM is Nvidia selling the answer to a problem that is currently inflating its own quotes.