Search

The real economics of AI compute: power, wafers, and the illusion of infinite scale

Yole Group’s latest analysis following the early November 2025 announcements confirms that the AI boom is real, but the timing isn’t. Silicon is ready months before the gigawatts.

Over the last 18 months, the semiconductor industry has been living through a generative-AI gold rush. Capital expenditure commitments, GPU shipments, and data-center power projections are growing at rates the market hasn’t seen since the dot-com or crypto-mining booms.

The latest Nvidia GTC keynote made the scale explicit, and it’s time to separate what’s physically possible from what’s speculative.

The GTC shockwave: 4 million Hoppers, 3 million Blackwells, and 7 million more IN THE PIPELINE

Nvidia’s GTC Washington D.C. 2025 keynote confirmed what many suspected: the company’s GPU engine is running at industrial scale.

  • 4 million Hopper GPUs (H100/H200 class) shipped across 2023-2025.
  • 3 million Blackwell GPUs already produced between Q1-2024 and Q3-2025.
  • A pipeline of 7 million more Blackwell and Rubin GPUs to be shipped by the end of 2026.

At the same time, Nvidia teased a total of 20 million “GPU dies” for 2025-2026, but this number is misleading. It refers to dies, not packages. Since each data-center GPU package integrates multiple dies, the physical shipment count in packages is roughly half that figure. In other words, the 20 million figure makes the ramp-up sound even larger than it is.

Translating hype into hardware: what 7 million Blackwell + Rubin really means

When we translate Nvidia’s statements into system-level terms, the scale becomes easier to grasp.

hugo-antoine-pro
Hugo Antoine Technology & Market Analyst, Computing and Software at Yole Group
Seven million GPU packages correspond roughly to 97,000 racks of compute systems such as NVL72 or NVL144. Assuming a PUE of 1.3, this implies around 18 GW of new Nvidia AI data center power capacity to be delivered by the end of 2026.

For context, OpenAI is for now targeting about 26 GW of total AI compute over the next four years, roughly 6.5 GW per year, split between 10 GW of Nvidia GPUs, 10 GW of Broadcom-custom AI ASICs, and 6 GW of AMD GPUs.  Some of these announcements likely overlap with the Stargate project: the 10 GW, $500 billion AI infrastructure program with Oracle and Softbank. According to an official letter from OpenAI’s Chief Global Affairs Officer, roughly 7 GW of that capacity is already planned for deployment over the next three years. But of course, it doesn’t stop there… OpenAI has also committed to renting $250 billion worth of compute from Microsoft Azure and $38 billion from AWS, and has signed multi-billion-dollar deals with CoreWeave.

And that’s just OpenAI. The rest of the hyperscaler, neocloud, and model-training ecosystem, including xAI, Anthropic, Mistral, CoreWeave, Google, AWS, Microsoft, Meta, Oracle, and others, is stacked on top of this demand.

OpenAI suggests that the US should double its annual power deployment from roughly 50 GW today to around 100 GW per year to sustain the next wave of AI infrastructure growth.

The numbers are unprecedented, and the physics are starting to bite.

A supply chain stretched to Its limits: wafer and power reality check

The semiconductor industry loves volume ramps, but can Logic wafers, CoWoS packaging, HBM memory, and Power really scale as quickly as the latest announcements suggest?

This article is reserved for logged-in users. You have 80% left to discover.
up