Major AI vendors are moving away from flat-rate per-seat subscriptions toward token consumption and outcome-based pricing models, according to The Next Web. The shift marks the end of the adoption-phase loss leaders that locked enterprises into AI tooling at predictable costs. “We believe software value should align directly with customer success, not headcount,” Zendesk’s president for products, engineering and AI, Shashi Upadhyay, told TNW. For enterprises running thousands of AI queries daily, the result is higher and less predictable bills.
The Local Hardware Response
Enterprises and knowledge workers are responding by moving routine AI workloads off the cloud entirely. AI PCs with neural processing units can now handle basic and mid-level generative tasks (summarization, drafting, code completion, data extraction) locally, with zero marginal cost per query after the hardware purchase, according to TNW. Consumers have been buying Mac Minis to run OpenClaw agent instances locally, avoiding per-query cloud costs. For heavy users, the payback period on a $1,500 AI PC is months, not years.
The DRAM crisis has pushed memory costs higher, making AI PCs more expensive to buy, but the cost-per-query advantage holds for high-volume, low-complexity tasks, according to TechRadar.
Cloud Is Not Going Away
The shift is not cloud versus local. Training frontier models, running complex multi-step agents, and processing enterprise-scale data still require cloud infrastructure. Alphabet raised its capex guidance to $205 billion this year as Google Cloud revenue jumped 82%, according to TNW. The hyperscalers are building for growing cloud AI demand. But consumption pricing gives enterprises a financial incentive to segment workloads: keep complex, high-value tasks on cloud infrastructure and move every routine query that can run locally off the meter.
The Agent Cost Architecture
For teams running autonomous agents, the economics create a two-tier deployment architecture. Agents handling routine, high-frequency tasks (email triage, document summarization, code review, data extraction) become candidates for local deployment on hardware with zero per-query costs. Agents requiring frontier model capabilities, multi-step reasoning, or access to large context windows remain cloud-bound. The pricing shift makes this segmentation not just technically sensible but financially necessary.