Share with your CTO
Microsoft is betting on AMD’s rack-scale Helios system to carry Azure’s next wave of large-model inference, introducing the ND MI455X v7 virtual machine built on a 72-GPU, 31TB HBM4 rack that integrates AMD’s Instinct MI455X accelerators, EPYC Venice CPUs, and Pensando networking into a single unified design. AMD begins shipping in the second half of 2026. Two new CPU-only VM families, HDv2 and HXv2, carve Venice into separate roles for AI data pipelines and semiconductor simulation respectively, making this an infrastructure-layer expansion across compute, networking, and cooling, not just a GPU procurement swap.
What this means for your business
The jump from eight-GPU VM slices to a 72-GPU rack-scale unit is a structural shift in how Azure prices and delivers inference capacity. Organizations currently planning inference deployments on ND MI300X v5, Azure’s existing AMD GPU VM, are evaluating a product whose successor hasn’t disclosed per-VM GPU count, memory allocation, pricing, or regional availability. If your roadmap assumes MI300X v5 pricing as a baseline for 2027 cost models, that baseline is now provisional.
The deeper play here isn’t the GPU count, it’s the vertical integration. AMD is following the pattern that gave Google’s TPU pods and AWS’s Trainium clusters their cost advantage: when a single vendor controls the accelerator, the CPU, the network card, and the cooling architecture, the hyperscaler can negotiate the whole stack rather than assembling it from parts. Microsoft embedding Pensando DPUs alongside MI455X and EPYC Venice means AMD’s footprint in Azure now spans the data path from storage offload through inference compute, which is exactly the kind of supplier entrenchment that makes future switching expensive. CTOs at enterprises running multi-cloud inference strategies should treat this as evidence that Azure’s AMD stack is becoming a coherent platform, not a second-source hedge against Nvidia.
AMD’s headline figure of 2.9 exaFLOPS at FP4, a low-precision numeric format that reduces memory load per calculation, is a design-ceiling number measured under ideal conditions with no contention, no network latency, and no software overhead. Real inference throughput for a specific model at a specific batch size will be materially lower, and ROCm’s (AMD’s GPU software layer, comparable to Nvidia’s CUDA) maturity at rack scale remains unproven in production. The decision this actually reframes is whether your organization’s Nvidia commitments deserve a renegotiation posture: not because AMD wins on paper specs, but because a credible Azure-native AMD rack changes the negotiating dynamic on Nvidia reserved capacity renewals happening now.
Concept deep-dive: Rack-scale design
Traditional cloud GPU infrastructure bolts accelerator cards into general-purpose servers, then connects those servers with standard networking. Rack-scale design, by contrast, engineers the accelerators, CPUs, interconnects, power delivery, and liquid cooling as a single system from the start, the way a custom racing engine is built around a chassis rather than dropped into one. The business consequence is higher sustained throughput per watt and lower latency between GPUs, but it also means the vendor controls the full stack, which concentrates both performance gains and switching costs.
Based on reporting from Microsoft Adopts AMD “Helios” for Azure, Extending AI Inference Infrastructure to 72-GPU Racks, originally published 2026-07-20 21:48:00.

