Business

Curb Runaway Agentic Coding Costs

Just months after many organizations incentivized “tokenmaxxing” through AI consumption leaderboards, many pumped the brakes due to exploding costs. Experts tell how enterprises can control AI token expenses without slowing innovation.
  • Article:Business
  • Key Play:Enterprise AI
  • Nutanix-Newsroom:Article

September 2, 2026

Months after companies celebrated employees who burned the most AI tokens, the same organizations are scrambling to contain the bill.

AI adoption can be divided into four overlapping “waves,” according to Paul Updike, a senior director of technical marketing engineering at Nutanix. He broke them into knowledge (generative AI tools, including retrieval augmented generation, or RAG); agentic coding; agentic applications; and eventually, an agentic “free-for-all” where AI tools have broad permissions to take a wide range of actions across the enterprise. 

For now, agentic coding is the “killer app” driving the most change in many organizations and is also the greatest expense. Eventually, Updike said, AI may push organizations to implement greater control over how applications are deployed than they have since the rise of the web in the 1990s. But at the moment, leaders are grappling with a more pressing problem: runaway costs.

Agentic coding uses AI agents to autonomously write, test and debug software across multiple steps, consuming far more tokens than standard chatbot interactions and driving up enterprise AI costs.

According to some estimates, 40% of companies now spend at least $10 million a year on AI.

“Agentic coding is causing a lot of budget pressure,” Updike said. 

“We’re seeing companies say: ‘What did we do? This costs way more than before.’ But there’s always pressure to adopt a new, better model, and that model is always going to use exponentially more tokens.”

RELATED Neoclouds and Agentic AI Shift Focus to CPUs
In this video interview, AMD's Brayden Mahdavi says the shift from training to inference is reshaping every layer of the stack.
  • Nutanix-Newsroom:Article, Video
  • Products:Nutanix Cloud Infrastructure (NCI), Nutanix Kubernetes Platform (NKP), Nutanix Unified Storage (NUS)

August 18, 2026

In interviews with The Forecast, Updike and other experts shared their ideas for getting AI costs under control without hindering innovation. He said many are realizing that that on-premises infrastructure can pay for itself fairly quickly, rather than incurring recurring token costs to a cloud provider. But supply chain constraints can hamper progress.

Why Move AI Workloads On-Premises?

Anthony Jackman is chief innovation officer for the data center and cloud services provider Expedient. He sees organizations increasingly taking a critical look at their cloud environments and moving resources that aren’t a fit.

“We don’t have the view that literally everything should run with us,” Jackman said. “We see [customers] peeling back the things that were mass-migrated to the hyperscalers that weren’t a great fit. We call it cloud rebalancing. There are workloads that are genuinely a very good fit, and we encourage our clients to put every workload in the best place for that workload.”

RELATED Own Your Token Machine
In this video interview, Liqid Founder and CTO Sumit Puri argues the path to affordable AI inferencing runs through composable IT infrastructure that dynamically taps into pools of GPUs, CPUs and DRAM. This helps optimize usage to manage tokenomics.
  • Nutanix-Newsroom:Article, Video
  • Products:Nutanix Kubernetes Platform (NKP)
  • Use Cases:AI ML, Cloud Native

August 29, 2026

For AI workloads, Updike said, on-premises infrastructure increasingly seems like the answer to that question, at least over the long run. He noted that Nutanix itself could hypothetically “make money” by bringing AI on-premises, with the upfront capital expense paying for itself in around two years.

“The only thing keeping that from happening right now is supply chain, and the ability to buy the gear,” Updike said.

Why Turn to Neoclouds (at Least for Now)

Neoclouds offer economical, ready access to scarce GPU compute for organizations not yet ready to build on-premises AI environments.

Neoclouds are AI-first cloud providers initially built to support large-scale GPU compute. They offer an alternative to hyperscalers for organizations that want an alternative to hyperscalers but aren’t yet ready to invest in massive on-premises AI environments. In addition to providing ready access to scarce resources, neoclouds can be more economical than hyperscaler environments for sustained GPU use, explained Updike.

RELATED The Rapid Rise and Future of Neoclouds
Industry experts explore whether specialized clouds for AI infrastructure are just a fad or if they will stand the test of time.
  • Article:Technology
  • Key Play:Enterprise AI
  • Nutanix-Newsroom:Article

May 22, 2026

“The neocloud providers made a gamble to buy the hardware, and now they can bring customers on-premises with them,” Updike said. “They’re a place where you can take an intermediary step if your goal is to bring AI on premises.”

Why It’s Important to Enforce Strategic Model Selection

Routing routine tasks to lower-cost models rather than frontier models can sharply reduce token consumption without sacrificing results.

Within many organizations, employees are currently using advanced AI models for relatively simple tasks, needlessly driving up token consumption and costs.

Juan Orlandini, chief technology officer at global solutions integrator Insight, admitted with a laugh that he recently used a frontier model to check the next day’s weather forecast. 

“That’s a gazillion token dollars,” Orlandini said. “I could very well have brought up the weather interface on my computer or looked at my watch.”

RELATED Agent Gateway Enforces AI Token and Traffic Control
AI agents are multiplying and IT governance will need to keep pace in order to manage chaos, cost and data security. Nutanix experts describe a new kind of gatekeeper stepping into the breach, bringing speed to builders and control to enterprise IT teams.
  • Article:Technology
  • Key Play:Enterprise AI
  • Nutanix-Newsroom:Article

June 25, 2026

Ashwini Vasanth, group product manager for Nutanix Enterprise AI technology, noted that features such as the Nutanix Agentic Gateway sit between an organization’s users and AI models, enabling IT to control, monitor, and route all AI traffic. Vasanth noted that organizations can experiment with sending users to lower-cost models to determine which provides the best balance of usability and cost control.

“Let’s say today, there’s a new open-source model that got released,” Vasanth said. “I want to maybe send 20% of my traffic there and see if people complain. If they don’t, then I know it’s okay, and I can switch over and end up saving costs in the long run.”

Why It’s Critical to Control Consumption

As vendors switch to pay-per-use pricing, organizations are capping individual AI spending to keep runaway token costs in check.

The idea of limiting AI use would have been nearly unthinkable within many organizations at the beginning of this year, when many were incentivizing consumption with “tokenmaxxing” leaderboards. But as AI vendors began switching from flat-fee to pay-per-use pricing, some organizations are now capping consumption. One notable example is Uber, which abruptly began capping employees at $1,500 per month in AI spending after previously ranking teams by AI usage.

“The metered model is starting to expose the true cost of tokens,” said Orlandini. “So you have to be very intelligent about the tokenomics. Part of the tokenomics conversation is how do I meter, and make sure that my tokens are being used wisely and appropriately?”

RELATED Token Era Forces Mastery of AI Economics
Four decades after mainframe engineers optimized every MIPS, a new generation is learning the same lesson with AI tokens. The currency changed but the discipline requires hasn’t, says Daryush Ashjari, Nutanix CTO for Asia Pacific and Japan.
  • Article:News
  • Key Play:Enterprise AI
  • Nutanix-Newsroom:Article

July 21, 2026

Jackman advises taking control. 

“Unfettered access to AI, especially by people that don't have that much experience with it yet, can cause costs to spiral out of control very, very quickly.”

Updike noted that software engineers have a history of working long hours on their laptops to push projects forward. Today, those extra hours may come with an unacceptably high AI bill.

“In the past, I would just stay up all night working on something,” Updike said. “If you do that with AI, and you’re just hitting ‘yes’ and ‘go,’ you’re costing the company a fortune.”

“If you’re putting in a 20-hour workday with an AI model, you’d better have a reason,” Updike added. “Companies are starting to ask: How much of this incredible technology are you going to use today?”

Related:

Editor’s note: Learn about Nutanix Enterprise AI (NAI) 2.8, and the upcoming general availability of Nutanix Kubernetes Platform (NKP) 2.19, along with new incentives, programs, and resources designed to help partners accelerate growth on emerging AI opportunities in this press release: Nutanix Gives Enterprises the Freedom to Run Production Agentic AI Their Way.

Calvin Hennick is a technology journalist covering enterprise IT, cloud and AI infrastructure. Connect with him on LinkedIn.

© 2026 Nutanix, Inc. All rights reserved. For additional information and important legal disclaimers, please go here.

Key Takeaways:

  • Agentic coding is the "killer app" driving AI adoption but also the greatest expense, putting budget pressure on organizations, according to Paul Updike of Nutanix.
  • On-premises infrastructure can pay for itself in about two years, but supply-chain constraints limit near-term adoption, Updike said.
  • Routing routine tasks to lower-cost models and capping consumption are immediate ways to rein in token costs, said Juan Orlandini of Insight.

Related Articles