Agentic AI tasks consume roughly 1,000 times more tokens than single-turn code reasoning or chat interactions, according to McKinsey. AMD used that fact as the opening argument at its Advancing AI 2026 conference in San Francisco on July 23, positioning its full hardware portfolio around a single thesis: whoever reduces token cost per task wins enterprise agent deployments.
The Token Economics Problem
The cost gap between chatbots and agents is structural. Each tool call, each reasoning step, each verification loop compounds the bill. McKinsey’s analysis found that about 60% of an agentic task’s costs come from refining, checking, and reverifying answers rather than generating the initial response, according to IT Pro.
Enterprise adoption is accelerating despite the costs. In a survey of over 1,000 IT decision makers, DigitalOcean found that 53% reported productivity and time savings from AI agents, while 44% said agents had created new business opportunities.
AMD’s Hardware Response
AMD’s answer is the MI455X GPU, announced during CEO Lisa Su’s keynote. The company claims up to 18x lower token cost than its predecessor, the MI355X, with up to 34x higher token throughput, according to IT Pro. AMD’s broader Helios rackscale system, which pairs 72 MI455X GPUs with 18 6th Gen EPYC “Venice” CPUs, delivers up to 30% more inference tokens per dollar than competing solutions, per AMD’s press release.
“The cost of tokens as the world moves to inference is so important,” AMD executive Jack Dieckman told IT Pro.
A Three-Tier Data Center for Agents
Beyond GPUs, AMD outlined a new data center architecture built around three distinct CPU tiers. Agent sandbox CPUs run and coordinate agents, execute spawned code, and handle I/O. AI host node CPUs feed GPUs across rack-scale infrastructure. General-purpose CPUs run databases, storage, and application services that agents interact with.
The EPYC 9006 SP7, the agent sandbox tier, scales to 256 cores and 512 threads for high-density agent execution. AMD’s pitch: more agents per watt, per dollar, per rack.
The Hybrid Bet
Michael Nordquist, CVP of product marketing at AMD, acknowledged the tension between local and cloud-hosted inference. “It’s not all one way or all the other, it’s going to be hybrid,” he told IT Pro. “For some things local models work great, for some things you’re going to want that frontier model available for agent behaviour.”
That hybrid positioning lets AMD sell into both camps: enterprises running agents on local infrastructure to control costs, and cloud providers hosting frontier agent workloads at scale. Launch customers for Helios include Anthropic (deploying up to 2 gigawatts of MI455X GPUs), OpenAI, Meta, Microsoft, and Oracle, according to AMD.