Months after companies celebrated employees who burned the most AI tokens, the same organizations are scrambling to contain the bill.
AI adoption can be divided into four overlapping “waves,” according to Paul Updike, a senior director of technical marketing engineering at Nutanix. He broke them into knowledge (generative AI tools, including retrieval augmented generation, or RAG); agentic coding; agentic applications; and eventually, an agentic “free-for-all” where AI tools have broad permissions to take a wide range of actions across the enterprise.
For now, agentic coding is the “killer app” driving the most change in many organizations and is also the greatest expense. Eventually, Updike said, AI may push organizations to implement greater control over how applications are deployed than they have since the rise of the web in the 1990s. But at the moment, leaders are grappling with a more pressing problem: runaway costs.
Agentic coding uses AI agents to autonomously write, test and debug software across multiple steps, consuming far more tokens than standard chatbot interactions and driving up enterprise AI costs.
According to some estimates, 40% of companies now spend at least $10 million a year on AI.
“Agentic coding is causing a lot of budget pressure,” Updike said.
“We’re seeing companies say: ‘What did we do? This costs way more than before.’ But there’s always pressure to adopt a new, better model, and that model is always going to use exponentially more tokens.”
In interviews with The Forecast, Updike and other experts shared their ideas for getting AI costs under control without hindering innovation. He said many are realizing that that on-premises infrastructure can pay for itself fairly quickly, rather than incurring recurring token costs to a cloud provider. But supply chain constraints can hamper progress.
Anthony Jackman is chief innovation officer for the data center and cloud services provider Expedient. He sees organizations increasingly taking a critical look at their cloud environments and moving resources that aren’t a fit.
“We don’t have the view that literally everything should run with us,” Jackman said. “We see [customers] peeling back the things that were mass-migrated to the hyperscalers that weren’t a great fit. We call it cloud rebalancing. There are workloads that are genuinely a very good fit, and we encourage our clients to put every workload in the best place for that workload.”
For AI workloads, Updike said, on-premises infrastructure increasingly seems like the answer to that question, at least over the long run. He noted that Nutanix itself could hypothetically “make money” by bringing AI on-premises, with the upfront capital expense paying for itself in around two years.
“The only thing keeping that from happening right now is supply chain, and the ability to buy the gear,” Updike said.
Neoclouds offer economical, ready access to scarce GPU compute for organizations not yet ready to build on-premises AI environments.
Neoclouds are AI-first cloud providers initially built to support large-scale GPU compute. They offer an alternative to hyperscalers for organizations that want an alternative to hyperscalers but aren’t yet ready to invest in massive on-premises AI environments. In addition to providing ready access to scarce resources, neoclouds can be more economical than hyperscaler environments for sustained GPU use, explained Updike.
“The neocloud providers made a gamble to buy the hardware, and now they can bring customers on-premises with them,” Updike said. “They’re a place where you can take an intermediary step if your goal is to bring AI on premises.”
Routing routine tasks to lower-cost models rather than frontier models can sharply reduce token consumption without sacrificing results.
Within many organizations, employees are currently using advanced AI models for relatively simple tasks, needlessly driving up token consumption and costs.
Juan Orlandini, chief technology officer at global solutions integrator Insight, admitted with a laugh that he recently used a frontier model to check the next day’s weather forecast.
“That’s a gazillion token dollars,” Orlandini said. “I could very well have brought up the weather interface on my computer or looked at my watch.”
Ashwini Vasanth, group product manager for Nutanix Enterprise AI technology, noted that features such as the Nutanix Agentic Gateway sit between an organization’s users and AI models, enabling IT to control, monitor, and route all AI traffic. Vasanth noted that organizations can experiment with sending users to lower-cost models to determine which provides the best balance of usability and cost control.
“Let’s say today, there’s a new open-source model that got released,” Vasanth said. “I want to maybe send 20% of my traffic there and see if people complain. If they don’t, then I know it’s okay, and I can switch over and end up saving costs in the long run.”
As vendors switch to pay-per-use pricing, organizations are capping individual AI spending to keep runaway token costs in check.
The idea of limiting AI use would have been nearly unthinkable within many organizations at the beginning of this year, when many were incentivizing consumption with “tokenmaxxing” leaderboards. But as AI vendors began switching from flat-fee to pay-per-use pricing, some organizations are now capping consumption. One notable example is Uber, which abruptly began capping employees at $1,500 per month in AI spending after previously ranking teams by AI usage.
“The metered model is starting to expose the true cost of tokens,” said Orlandini. “So you have to be very intelligent about the tokenomics. Part of the tokenomics conversation is how do I meter, and make sure that my tokens are being used wisely and appropriately?”
Jackman advises taking control.
“Unfettered access to AI, especially by people that don't have that much experience with it yet, can cause costs to spiral out of control very, very quickly.”
Updike noted that software engineers have a history of working long hours on their laptops to push projects forward. Today, those extra hours may come with an unacceptably high AI bill.
“In the past, I would just stay up all night working on something,” Updike said. “If you do that with AI, and you’re just hitting ‘yes’ and ‘go,’ you’re costing the company a fortune.”
“If you’re putting in a 20-hour workday with an AI model, you’d better have a reason,” Updike added. “Companies are starting to ask: How much of this incredible technology are you going to use today?”
Related:
Editor’s note: Learn about Nutanix Enterprise AI (NAI) 2.8, and the upcoming general availability of Nutanix Kubernetes Platform (NKP) 2.19, along with new incentives, programs, and resources designed to help partners accelerate growth on emerging AI opportunities in this press release: Nutanix Gives Enterprises the Freedom to Run Production Agentic AI Their Way.
Calvin Hennick is a technology journalist covering enterprise IT, cloud and AI infrastructure. Connect with him on LinkedIn.
© 2026 Nutanix, Inc. All rights reserved. For additional information and important legal disclaimers, please go here.