Microsoft Expands Azure with AMD Helios AI Racks
Microsoft will deploy AMD's Helios rack-scale platform on Azure in the second half of 2026, adding MI455X GPUs, EPYC Venice CPUs, Pensando networking and ROCm.
AMD Helios AI racks will add a new full-rack AI platform to the company’s cloud infrastructure in the second half of 2026. AMD says Microsoft will deploy Helios at scale for frontier-model inference, Azure AI services and customer workloads, expanding a partnership that already spans processors and cloud instances.
The announcement is significant because Helios is not a standalone accelerator. It combines AMD Instinct MI455X GPUs, sixth-generation EPYC “Venice” CPUs, Pensando networking and ROCm software as an integrated rack-scale system. Microsoft is therefore adopting AMD across the compute, networking and software layers needed to run large AI clusters.
What Microsoft is deploying
AMD says Helios shipments to customers, including Microsoft, will begin in the second half of 2026. Azure plans to use the system for inference across frontier models, its own AI services and applications operated by cloud customers.
The MI455X GPU is paired with EPYC Venice CPUs and AMD Pensando networking. ROCm provides the software environment, while the rack follows open industry standards intended to give cloud operators more flexibility in how systems are integrated and managed.
The agreement does not disclose how many racks or GPUs Microsoft has committed to deploy, nor does it provide pricing. “At scale” indicates a production infrastructure plan, but it should not be translated into an unsupported unit count.
Two new EPYC-powered Azure VM families
The expanded partnership also covers two Azure virtual-machine series powered by sixth-generation EPYC processors. Azure HDv2 is aimed at agentic AI and data pipelines, while HXv2 targets semiconductor design and other engineering workloads.
Those CPU-focused instances matter because AI infrastructure involves more than GPU tensor operations. Data preparation, search, orchestration and reinforcement-learning workflows can place heavy demands on general-purpose compute. Microsoft CEO Satya Nadella said customers want infrastructure optimized across that wider range of workloads.
Networking becomes part of the deal
Microsoft will also broaden its use of AMD Pensando data-processing units and integrate AMD technology with Azure Boost. DPUs offload networking and infrastructure tasks that would otherwise consume CPU resources, helping large clusters process connections more efficiently.
This full-stack approach is the main strategic point. A fast accelerator can still be constrained by data movement, networking and software maturity. Helios gives AMD a way to compete as a rack-level platform rather than as a GPU vendor supplying isolated components.
More choice does not mean Nvidia is out
The announcement expands Azure’s hardware portfolio; it does not say Microsoft is removing Nvidia infrastructure. Hyperscale cloud providers routinely use multiple processor architectures to match workload, availability and cost requirements.
Microsoft has continued to support Nvidia hardware across cloud and local AI use cases. ExstarHub previously covered how Microsoft enabled additional local AI features on Nvidia GPUs. Adding Helios creates another production option rather than a clean replacement.
That distinction also matters for developers. Nvidia’s CUDA ecosystem remains deeply established, while AMD is investing in ROCm compatibility and deployment tools. Microsoft putting Helios behind Azure services can reduce some of the integration burden for customers who consume managed compute rather than build racks themselves.
Why inference is the first headline workload
Training creates new model weights, while inference uses trained models to answer requests. At cloud scale, inference must balance latency, throughput, memory capacity and operating cost across sustained production traffic.
Microsoft identifies frontier-model inference as a central Helios workload. That gives AMD an opportunity to prove performance and reliability under the repeated, high-volume demand created by Azure AI services and customer applications.
The deployment also arrives as AI companies seek enormous pools of computing capacity from infrastructure partners. Deals such as the reported Meta and Anthropic compute discussions illustrate why hyperscalers want more than one path for bringing new capacity online.
What remains to be proven
The announcement establishes the partnership and expected shipping window, but commercial success will depend on deployment timing, ROCm maturity, utilization and performance in real Azure workloads. Microsoft and AMD have not published comparative cost or benchmark results for this rollout.
Helios is therefore a credible new infrastructure option, not proof that the accelerator market has already been reshaped. The important milestone is that Microsoft is preparing to operate AMD’s complete rack-scale AI stack in production.
Key takeaways
- Microsoft will deploy AMD Helios on Azure beginning in the second half of 2026.
- Helios combines MI455X GPUs, EPYC Venice CPUs, Pensando networking and ROCm software.
- Azure will use the platform for frontier-model inference, AI services and customer workloads.
- New HDv2 and HXv2 VM families will use sixth-generation EPYC processors.
- The expansion adds choice to Azure; it does not mean Microsoft is abandoning Nvidia.
Source: techpowerup.com
