Articles & Research

Briefing

Silicon for inference: Blackwell, Trillium, Ironwood and MI300

The accelerator roadmap that reshaped data-center design between 2023 and 2026.

Priya Raghunathan FACS

Chair, ACS Specialist Group on Semiconductors

February 2026 · 7 min read

Technicians in cleanroom suits handling a silicon wafer under lithography lighting
Technicians in cleanroom suits handling a silicon wafer under lithography lighting

Three vendors published three different bets on what the AI workload of 2026 would look like. All three converged on the same conclusion: inference, not training, would set the shape of the rack.

The roadmap, dated

Google announced TPU v5e general availability alongside A3 instances on 29 August 2023, and followed with TPU v5p and the AI Hypercomputer architecture on 6 December 2023. Trillium, its sixth-generation TPU, was announced on 14 May 2024 and reached general availability on 11 December 2024. Ironwood, the seventh generation, was unveiled on 9 April 2025 and explicitly framed by Google as a chip 'for the age of inference'; it reached general availability on 6 November 2025.

AMD launched the Instinct MI300 series on 6 December 2023, claiming a roughly eightfold generational performance gain in combination with ROCm 6, and on 2 June 2024 published an expanded roadmap: MI325X with 288GB of HBM3E for the fourth quarter of 2024, and the CDNA 4-based MI350 series for 2025 with a claimed 35x generational increase in AI performance.

NVIDIA unveiled the Blackwell platform, including B200 and GB200, at GTC on 18 March 2024, claiming up to a 25-fold reduction in cost and energy for large language model inference relative to Hopper. On 5 January 2026 the company announced the Rubin platform, pairing a next-generation GPU with the Vera CPU and NVLink 6.

“Memory capacity grew faster than compute for two generations, because inference on long contexts is bound by bandwidth, not arithmetic.”

Why the rack changed

Vendor performance claims should be read as marketing until independently reproduced, but the physical consequences of the roadmap are not in dispute. Memory capacity per accelerator grew faster than compute for two generations, because inference on long contexts is bound by memory bandwidth and capacity rather than by arithmetic. Rack power density moved from a range most enterprise facilities were built for to one most are not, forcing liquid cooling from an exotic option into a default assumption for new builds.

For members designing facilities, the operative planning number is no longer floor area. It is kilowatts per rack and the water or air path that removes them, decided years before the silicon that will occupy the space has been specified.

Professional implications

Two disciplines that used to sit far apart — model architecture and mechanical engineering — now constrain each other directly. A decision to serve a 2-million-token context window is a decision about HBM capacity, which is a decision about accelerator count, which is a decision about cooling.

ACS accreditation panels have begun asking computing programs how they teach that coupling. Students who can reason about attention mechanisms but not about the memory hierarchy they run on are trained for a job that no longer exists in isolation.

SemiconductorsArtificial intelligenceGreen computing

Join the professional body behind this work

ACS members receive our research first, free CPD and ethics modules every year, and a route to professional registration assessed by their peers.

Become a member