Tech firm AMD is urging enterprises to rethink their computing infrastructure as agentic AI adoption drives up token consumption and cloud computing costs, with the chipmaker pointing to distributed AI architectures as one way to manage expenses.
As autonomous AI agents move beyond pilot projects and become integrated into departmental workflows, their continuous operation could significantly increase usage-based costs.
“As AI moves from occasional prompts to continuous workloads via agentic AI, token consumption can grow quickly, where systems repeatedly reason, call tools, and iterate,” said Alexey Navolokin, Asia Pacific general manager at AMD.
“Each step consumes compute and system overhead [or tokens], and in cloud environments, usage-based costs can accumulate,” he added.
The growing costs come as enterprises increasingly deploy AI agents for more complex tasks. According to Anthropic’s State of AI Agents 2026 report, 57% of surveyed organizations already deploy agents for multi-stage workflows.
Navolokin said organizations are adopting agents partly because they can augment employees rather than simply automate jobs.
“AI agents can multiply what employees are able to accomplish, rather than simply replace human work. Agents can operate continuously, handle repetitive execution and coordinate multiple steps at a scale and pace that would be difficult for an individual to sustain, while employees remain focused on decisions, creativity, domain expertise and accountability,” he said.
Agents can also be adapted to different tasks. Small businesses, for example, can use them to support small teams, while creators can delegate scheduling and content distribution and professionals can use agents for research and information synthesis.
Within AMD, Navolokin said AI agents are already being used in semiconductor design to accelerate the identification and resolution of issues.
Moving AI workloads beyond the cloud
As the number of agents and their workloads increase, AMD said enterprises may need to reconsider relying exclusively on cloud infrastructure.
“AMD’s view is that agentic AI is a distributed infrastructure challenge, with the right mix of compute needed across cloud, data center, edge and AI PCs,” Navolokin said.
“A distributed AI architecture is one that supports AI across multiple environments – from cloud and data centers to edge systems and AI-enabled PCs. AI compute is becoming more distributed because no single environment is optimal for every workload.”
Navolokin said organizations need to practice “workload right-sizing,” or determining the most efficient and cost-effective infrastructure for particular AI workloads.
“Agentic AI can drive significant inference and token demand, so enterprises need to consider which workloads are most economical to run locally versus in the cloud. Frequent or latency-sensitive workloads may benefit from local or edge compute, while larger or highly elastic workloads can continue to use the cloud,” he said.
Cloud infrastructure would remain an important part of such an architecture, but AMD argues that running some workloads locally on AI PCs could reduce recurring inference expenses. While organizations would have to shoulder the upfront hardware cost, locally processed queries would not incur cloud token charges.
“The economics can become particularly compelling as AI usage grows,” Navolokin said.
AMD said its own analysis found that at a medium workload tier of about 5.7 million input tokens and 574,000 output tokens per user per day, a fleet of 500 AMD AI PCs operating under a 50% local and 50% cloud configuration could generate projected three-year savings of 40% to 60% compared with cloud-only deployments, depending on the cloud model used.
The company said the workload was intended to represent a knowledge worker actively using an agent harness such as Claude Code, Codex, or Hermes. According to AMD, a fully local deployment could produce higher savings, with the initial investment typically breaking even in less than 24 months.
AMD also estimated that an AI PRO R9700 desktop configuration could support about 18 million tokens of AI use per day, with electricity costing about $64.80 per month.
Under the company’s cloud comparison assumptions, AMD estimated a three-year cost of $6,533 for the desktop configuration compared with $81,108 for cloud-based usage.
The figures are based on AMD’s own analysis and assumptions and actual costs would depend on factors including workloads, cloud models, hardware utilization, electricity rates, and deployment configurations.
Navolokin said improvements in AI PC computing capabilities and smaller language models are also making it possible to process more AI workloads locally.
“AI PCs are increasingly being viewed as AI infrastructure investments because they bring AI compute closer to where employees, data and workflows actually reside,” he said.
“This is particularly relevant for agentic AI. Local systems with sufficient memory and compute can run models that support tool calling, context-rich workflows and multi-agent applications, allowing enterprises to reduce the amount of inference they need to send to cloud services for appropriate workloads. This can lower recurring cloud inference costs while also providing lower latency and greater control over where sensitive data is processed.”
Local processing could also play a role in AI governance as agents increasingly interact with enterprise applications and proprietary information, allowing organizations to keep some sensitive workloads on their own machines rather than sending data to external cloud services.
Infrastructure challenges remain
Moving toward distributed AI architectures, however, requires organizations to rethink how their technology stacks are designed.
“The biggest challenge is balancing cost, data control, and infrastructure capability. Agentic AI can drive significant inference and token demand, so enterprises need to consider which workloads are most economical to run locally versus in the cloud,” Navolokin said.
“Local AI also requires the right infrastructure, including sufficient memory, CPU and GPU capacity, security and software support.”
The growth of agentic AI could also change enterprise capacity requirements because a single employee may eventually operate several AI agents simultaneously.
“So, infrastructure planning will also need to evolve. One employee managing multiple agents can create much more concurrent demand for compute, memory, data access and orchestration than a traditional AI assistant. That means enterprises should plan for balanced infrastructure across CPUs, GPUs, memory, networking and software rather than focusing on GPU performance alone,” Navolokin said.
He said enterprises should retain flexibility as agentic AI technologies and workloads continue to change.
“Enterprises should build for flexibility. An open software ecosystem can allow teams to develop locally, test at scale and deploy workloads where they make the most sense, while reducing lock-in as agentic AI evolves. The goal should be to create an environment where employees and AI agents can work together securely, efficiently and at scale.”
“At AMD, our approach is to provide an open ecosystem portfolio spanning data center, edge and AI PCs, supported by open software and standards that give customers flexibility as their AI deployments evolve,” Navolokin said.


