AMD will begin shipping Helios, its first rack-scale AI system, to Microsoft, Meta, OpenAI, Oracle, and Tata Consultancy Services later this year. CNBC reports that Helios is the first direct competitor to Nvidia’s Grace Blackwell and Vera Rubin systems, and Microsoft will deploy it in Azure data centers for frontier model inference.

AMD shares climbed more than 4% on the announcement. Microsoft stock rose over 1%.

What Helios Is

Helios integrates four components AMD builds in-house: GPUs, CPUs, networking, and software. AMD’s data center head Forrest Norrod told CNBC the system targets “the best total cost of ownership, the lowest cost per token, all in.” CEO Lisa Su said Helios has “significant benefits” over Nvidia’s rack-scale systems “when you’re talking about inference and when you’re talking about memory bandwidth and memory capabilities.”

The Futurum Group estimates Helios will cost between $5 million and $5.5 million per rack, compared to $3.5 million to $4 million for Nvidia’s Vera Rubin. At up to 7,000 pounds, Helios is wider and heavier than Nvidia’s offering.

The Customer List

Microsoft’s Azure deployment will power frontier model inference, AI services, and two new computing instances on AMD’s latest “Venice” CPUs: one for agentic AI and data pipelines, another for semiconductor design. “We are expanding the Azure infrastructure portfolio with AMD Helios to give customers the performance, scale and choice they need to build and run the next generation of AI applications,” CEO Satya Nadella said in a press release cited by CNBC.

Meta committed to up to 6 gigawatts of AMD GPUs in February, starting with 1 gigawatt deployed on Helios racks this year. OpenAI and Oracle made major deployment commitments. India’s largest IT company, Tata Consultancy Services, also signed on.

AMD says eight of the top 10 AI companies now run workloads on its Instinct GPUs, including Cohere and SpaceXAI.

The Market Gap

Nvidia controls more than 95% of the data center GPU market, according to the Futurum Group. AMD holds roughly 4.5%. Daniel Newman, analyst and CEO of the Futurum Group, told CNBC that AMD “can get to 20% and 25%. And by the way, this is hundreds of billions of dollars of revenue.”

AMD plans to book tens of billions in data center AI revenue starting in 2027, with the majority from Helios. Data centers already make up the majority of AMD’s revenue, up 57% year over year in Q1 2026.

Inference Economics for Agent Workloads

The agent infrastructure angle is specific. Microsoft is building dedicated Azure instances on Helios for “agentic AI and data pipelines,” which means the hardware competition directly targets the workloads that autonomous agents generate. Agent systems running continuous inference loops are the most cost-sensitive AI workloads: per-token cost determines whether an agent deployment is economically viable at scale.

AMD pricing Helios higher per rack ($5M+ vs Nvidia’s $3.5M+) but competing on cost-per-token and memory bandwidth suggests the system is optimized for the high-throughput, memory-intensive inference that agent workloads demand. If the total cost of ownership math works, enterprises running large agent fleets have a second vendor option for the first time.

Nvidia’s 95% market share meant agent inference costs were set by a single vendor. A credible second supplier, backed by commitment letters from Microsoft, Meta, and OpenAI, introduces price competition into the cost structure that every agent deployment depends on.