AMD's Helios bets AI networking on open Ethernet

  • AMD launches Helios, a 72-GPU rack built on UALink over Ethernet, Pensando DPUs and AI NICs rather than proprietary interconnects
  • Anthropic, OpenAI and Meta commit to gigawatt-scale Helios deployments; Microsoft and Oracle sign on
  • AMD claims 30% more tokens per dollar than Nvidia's Rubin NVL72, though key networking performance figures come from AMD simulations

AMD ADVANCING AI 2026, SAN FRANCISCO — AMD launched Helios, its first full rack-scale AI system, built around the new Instinct MI455X GPU — and around a networking argument aimed directly at Nvidia. Every fabric in the rack, from the front-end network to the GPU-to-GPU interconnect, runs on Ethernet and open standards rather than the proprietary links that dominate today's AI clusters.

Helios connects 72 MI455X GPUs into a single scale-up domain with 31 TB of unified HBM4 memory, 2.9 exaflops of peak FP4 compute and 260 TB/s of scale-up bandwidth, alongside AMD's 6th Gen EPYC "Venice" CPUs and Pensando networking silicon. AMD said the rack delivers up to 30% more tokens per dollar than Nvidia's Rubin NVL72 system and 50% more memory capacity. The system is in production now, with volume deployments ramping through the second half of 2026.

AMD is racking up some impressive customer wins. Anthropic committed Wednesday to deploy up to 2 gigawatts of MI450-series GPUs in Helios racks, with the first gigawatt beginning in the first half of 2027. AMD committed to invest up to $5 billion in Anthropic equity as part of the deal. That follows a 6-gigawatt multigenerational agreement with OpenAI, a 6-gigawatt custom GPU deal with Meta, Microsoft's commitment this week to deploy Helios at scale on Azure and Oracle's plan to deploy 50,000 MI450 GPUs on Oracle Cloud Infrastructure beginning in the third quarter of 2026.

"Between OpenAI, Meta, Anthropic, Microsoft, Oracle, as well as our strong neocloud partners, we expect several gigawatts of Helios infrastructure to be deployed over the coming year," said Andrew Dieckmann, AMD corporate VP and GM of the data center GPU business, in a briefing with press and analysts Wednesday.

Those figures deserve a caveat: the headline gigawatt numbers are "up to" commitments contingent on deployment milestones, and the Anthropic arrangement makes AMD both supplier and shareholder — part of a broader pattern of circular vendor financing propping up AI infrastructure deals across the industry. Anthropic, for its part, is hedging across silicon suppliers; the AMD capacity joins Nvidia GPUs, Amazon Trainium and Google TPU capacity in its compute portfolio.

Networking moves to the center of the AI stack

The most consequential part of the announcement for network operators and equipment vendors is the fabric architecture.

"At AI scale, networking is becoming as crucial as GPUs, CPUs, memory and storage to deliver continuous performance," said Soni Jiandani, AMD SVP and GM of the Networking Technology Business Unit.

(Editor's note: We told you so. We told you two years ago.)

Helios uses three distinct networks, all Ethernet-based. The front end runs on the third-generation Pensando Salina DPU at 400G, offloading software-defined networking, security, storage disaggregation and key value (KV) cache management for AI context. The scale-up fabric — the domain Nvidia serves with proprietary NVLink — uses the first generation of UALink over Ethernet (UALoE), connecting all 72 GPUs at 260 TB/s.

Scale-out runs on the second-generation Pensando Vulcano AI NIC at 800G, with three NICs per GPU delivering 2.4 Tbps of bandwidth per GPU, supporting Ultra Ethernet Consortium (UEC) transports.

A new Fabric Manager layer adds zero-touch provisioning, fabric-wide observability and "virtual pods" that carve the 72-GPU domain into isolated failure zones — network automation applied to the AI fabric itself.

An advantage to operators for the AMD approach is that Ethernet-based scale-up lets cloud providers and carriers apply four decades of Ethernet tooling and operational practice to AI clusters instead of adopting a vendor-specific interconnect. That continuity argument has been building since AMD shipped the industry's first UEC-ready NIC, the Pollara 400, and it echoes what Arista CEO Jayshree Ullal has called Ethernet's position as the "eventual winner" for AI networking). After surpassing Cisco in the data center switching market, Arista is leveraging Ethernet to set its sights on Nvidia

The market is moving toward Ethernet, which surpassed InfiniBand in AI back-end networks in 2025, according to Dell'Oro Group, and accounted for about two-thirds of data center switch sales in AI clusters in the first quarter of 2026. That's a striking reversal from 2023, when InfiniBand held roughly 80% of the market.

But Nvidia is hardly ceding the ground: its InfiniBand sales more than tripled in the quarter on the Blackwell Ultra ramp, and it pushes its own Spectrum-X Ethernet platform alongside NVLink.

Oracle Cloud Infrastructure has deployed Pensando DPUs across multiple generations for roughly a decade, converging network, storage and security functions onto the DPU. DPUs are foundational not just for networking but for security, storage and tenant isolation in OCI's infrastructure, said Rajagopal Subramaniyan, SVP of networking at Oracle Cloud Infrastructure, at Wednesday's briefing. One hyperscaler reclaimed 22 CPU cores per server by offloading its software load balancer to the DPU, Jiandani said; AMD's documentation identifies the customer as Microsoft, running the workload as part of Azure's Accelerated Connections service.

Vultr runs AMD networking and GPU infrastructure for frontier model customer. Larger models no longer fit on single systems, driving distributed inference that spreads workloads across many GPUs and makes the network fabric part of the computation, said David Gucker, COO and co-founder at Vultr.

Read the footnotes on the performance claims

AMD's networking numbers merit scrutiny. The claim that Vulcano's 50% scale-out bandwidth advantage translates to 13% faster LLM training time comes from AMD engineering modeling and synthetic benchmark simulation of an 8,000-GPU system, not from measured deployments. AMD's own endnotes say the results assume ideal behavior. A claimed 33% network switching cost savings on a 32,000-GPU cluster is likewise an estimate, based on Vulcano enabling 1.6T Tomahawk 6 switch platforms versus 800G alternatives. The rack-level compute and bandwidth figures, by contrast, are measured, and third-party firm Signal65 independently validated competitive throughput and better cost-per-token for the prior-generation MI355X against Nvidia's B200 on several open-source models.

Those economics matter because the buying pattern has flipped. "More and more purchasing decisions are made inferencing first," Dieckmann said. That shift favors whoever delivers the lowest cost per token at scale — and it foregrounds the network, which determines GPU utilization and latency in distributed inference.

Inference now accounts for the majority of AI compute — two-thirds in 2026, Deloitte estimates, up from one-third in 2023 — and clusters are increasingly purchased with inference as the primary workload from day one, Dieckmann said.

The CUDA moat gets a 2026 stress test

Historically, CUDA, , Nvidia's proprietary programming platform, has been a moat protecting Nvidia's market share, but that moat is vanishing, Dieckmann said. "I used to talk to our customers about CUDA a fair bit. I have almost zero conversations with our customers about CUDA at this point in time," he said. Customers program at higher abstraction levels and AI coding agents now handle much of the porting work: "Software obstacles are much, much easier to overcome for our customers now than they were in some years past."

That's ironic if true — AI built Nvidia's fortune and AI is now eroding the foundation.

AMD is driving its anti-CUDA push with ROCm.AI, an AI-driven development layer atop its open-source ROCm stack. ROCm.AI's Hyperloom component runs agentic optimization loops — profiling a model, finding bottlenecks, tuning and rewriting GPU kernels, validating gains — with no human in the loop. In a demonstration, the system improved inference performance on a MiniMax model by 38% autonomously. AMD publishes skills for Claude Code, Codex, Cursor and Gemini agents, and said release-to-release inference performance on frontier open models has improved 3.3x on average over roughly the past year through software alone. However, Ramine Roane, AMD corporate VP of product application engineering for the AI group, cautioned that point-in-time GPU benchmark comparisons circulating online are largely artifacts of which vendor optimized a given model first: "They're not comparing system to system."

Learn more about AI and data center architecture

'AI tailwind' keeps Arista at No. 1 spot in data center switching — now Nvidia looms

Nvidia eyes data center Ethernet as its next multi-billion-dollar biz