AMD launches full-stack AI compute for agentic era

  • AMD launches Helios racks, Instinct MI400 series GPUs and 6th Gen EPYC "Venice" CPUs as one integrated AI platform
  • Anthropic, OpenAI and Meta executives detail gigawatt-scale deployment plans, with Helios shipments beginning this quarter
  • AMD projects the server CPU market will top $200 billion by 2030 as agentic AI workloads demand density-optimized cores

AMD ADVANCING AI 2026, SAN FRANCISCO AMD launched an AI infrastructure stack Thursday comprising a Helios rack-scale system, containing Instinct MI400 series GPUs inside and 6th Gen EPYC "Venice" server CPUs, with executives from Anthropic, OpenAI and Meta appearing on the keynote stage to detail deployments, with first shipments to begin this quarter.

Helios is in full production, with shipments starting at the end of the third quarter and ramping into 2027, and demand for the MI450 series is running above AMD's expectations, CEO Lisa Su said in a keynote kicking off the AMD Advancing AI 2026 conference Thursday. The rack connects 72 MI455X GPUs with 31 TB of unified HBM4 memory and 2.9 exaflops of FP4 compute, and AMD claims up to 30% more tokens per dollar than Nvidia's Rubin NVL72. The system's all-Ethernet, open-standards networking architecture differentiates Helios from Nvidia's proprietary interconnects. Schneider Electric released a 246 kW reference design for deploying the racks.

Hefty customers support the technology. Anthropic's deal, announced Wednesday, covers up to 2 gigawatts of MI450-series GPUs plus up to $5 billion in AMD equity investment. OpenAI — which holds warrants for up to 10% of AMD stock under its 6-gigawatt agreement — has been running GPT-class workloads on Helios racks for three months. "We expect that we'll be deploying Helios at massive scale starting towards the end of this year and then accelerating towards 2027," said Sachin Katti, head of infrastructure at OpenAI, during the keynote.

Tom Brown, Anthropic co-founder and chief compute officer, described how the lab evaluated AMD hardware before committing: a single engineer connected a prior-generation MI355X rack to Claude, told the AI to bring up the machine and left for the weekend. "We ended up with a graph of the actual performance of our leading model on it, just going up and up and up over the weekend," Brown said. Su said AMD offered engineering help and Anthropic declined.

Venice CPUs target a new workload: the agent sandbox

AMD launched the 6th Gen EPYC "Venice" CPU family, sporting 203 billion transistors on TSMC's 2-nanometer process, up to 1.8x the performance of the prior generation, which Su called one of the largest generational gains in EPYC history. Venice targets a category AMD calls agent sandboxes — servers where AI agents execute code, call tools and query data. These workloads run on CPUs rather than GPUs, and AMD built a dedicated 256-core Venice variant for them, scaling to 512 threads per socket.

AMD projects the server CPU market will exceed $200 billion by 2030. "It's not just a GPU game anymore. It's CPUs and GPUs. If anything, I think CPUs have become at least as important, if not more," said Santosh Janardhan, head of infrastructure at Meta, during the keynote. Agent-generated code "still needs to run somewhere," he said.

Meta has MI350-series GPUs in production on ranking and recommendation systems and will deploy the MI450 "across the board" under its 6-gigawatt agreement, Janardhan said.

Venice is in full production with rollouts at every major server OEM and cloud provider beginning in the fourth quarter. Customer demand is the strongest AMD has seen for a new EPYC generation, Su said.

Enterprise AI: MI350P, Cisco and 'agents swarm'

For enterprises that can't rebuild their facilities around liquid-cooled AI racks, AMD launched the Instinct MI350P, an air-cooled GPU that fits existing enterprise servers and supports models up to 260 billion parameters on a single card, delivering up to 4.2x more tokens per second per dollar than the competition, according to AMD. The company's own IT organization, running open-weight models on EPYC and MI350P hardware with intelligent routing between local and frontier models, cut token costs 43% with up to 3x faster response times for local workloads, said Dan McNamara, AMD SVP and GM of compute and enterprise AI.

Cisco is building the management layer for that distributed model.

"Humans click, but agents swarm," said Jeetu Patel, president and COO at Cisco, describing how agentic AI replaces spiky, human-driven inference with persistent, always-on infrastructure demand. Cisco Cloud Control will manage fleets of AMD local-AI devices alongside data center and cloud inference, with observability for agent behavior and what Patel called tokenomics — monitoring and containing each agent's token consumption, including quarantining overly consumptive agents. The management stack is in early availability now and reaches general availability in the U.S. in early fall, Patel said.

AT&T, meanwhile, launched its OTel 2.0 open source telecom models trained on AMD hardware at the keynote.

AMD also announced a partnership with Cerebras for ultra-low-latency disaggregated inference, pairing Helios racks with Cerebras' wafer-scale engines — AMD's counter to Nvidia's roughly $20 billion December deal for low-latency inference vendor Groq. The companies expect the combined system to deliver up to 5x the throughput of speed-optimized alternatives; Cerebras plans to deploy Helios in its own data centers, with the joint offering available first through Cerebras Cloud in the second half of 2026.

The roadmap and the caveats

Su closed with a multi-year cadence: a new Helios-class system every year, with Instinct MI500 GPUs in 2027. The new GPUs will deliver the largest generational leap in Instinct history and will give AMD leadership in scale-up compute, Nvidia's strongest domain, Su predicted. These will be followed by MI600 GPUs, the "Zen 7"-based "Florence" EPYC generation in 2028 and the "Zen 8" EPYC "Ravenna" in development for 2030. The successor racks already have names: Helios 500 pairs MI500 GPUs with EPYC "Verano" CPUs and next-generation Pensando "Como" and "Monza" networking silicon, with Helios 600 to follow.

The projections deserve caveats. A year ago AMD sized the AI accelerator market at $500 billion by 2028; it now says $1.4 trillion by 2030, with AMD's addressable market approaching $2 trillion — vendor forecasts made amid an industry running on circular financing arrangements, including AMD's own Anthropic stake. The performance comparisons are AMD benchmarks, not independent tests.

Learn more about AMD Advancing AI 2026

AMD launches Helios AI rack on open Ethernet networking, stacks gigawatt deals with Anthropic, OpenAI and Microsoft

AT&T launches OTel 2.0 open source telecom AI models trained on AMD as token use tops 1 trillion a month

Schneider Electric and AMD release Helios reference design for 246 kW AI data center racks