Agentic AI Needs an Operating Model, Not Just an Application Stack

By Mike Berthiaume, Senior Director, Systems Engineering – Nutanix

Every major shift in enterprise infrastructure has forced IT to answer the same question: how do we operate this at scale without reinventing the wheel every time the underlying technology changes? Virtualization answered it for servers. Kubernetes container orchestration helped answer it for how we deploy and operate cloud-native applications. But Agentic AI introduces a fundamentally different set of operational challenges. 

AI isn't just another application you deploy and walk away from. Models and their capabilities are evolving at an extraordinary pace. Infrastructure costs swing wildly depending on which workloads land where. Data becomes part of the runtime instead of sitting quietly behind it. And governance questions multiply almost overnight.  

Managing all of that with the old application playbook doesn't scale. What IT needs is a way to deploy, secure, govern, observe, and update AI applications consistently, regardless of which model, compute infrastructure, or data source sits underneath. The model is one piece. Operating the whole system around it is the real job. 

Stop Sending Every Request to Your Most Expensive Model

Once you accept that AI is an enterprise workload, the economics start to matter in familiar ways. Think about how you manage virtual machines. You don’t run every VM on the most powerful hardware available just because you have it. Yet that's roughly how a lot of teams treat their AI requests today, routing every prompt to the biggest, most capable model regardless of what the task needs. 

A simple summarization request doesn't need the same horsepower as a complex reasoning problem, and a gateway that understands the difference can route intelligently based on the user, the application, the required response quality, latency needs, data sensitivity, and cost.  

That turns efficiency from an abstract talking point into something measurable, such as cost per request, token consumption, latency, and whether workloads are quietly landing on expensive resources they didn't need. At enterprise scale, small improvements in those numbers compound quickly. The same applies to tokens. Repeatedly sending unnecessary context, duplicating prompts, or failing to reuse information can create significant cost and latency at scale. Intelligent caching and context optimization can reduce that overhead without compromising the user experience. The bigger shift is architectural. Infrastructure should be optimizing itself around the outcome you're after, instead of leaving every application developer to make that call alone. 

Governance Has to Be Part of the Operating Model

Efficiency gets attention because it's easy to quantify. But for most IT executives, the bigger concern is trust. That leads directly to questions around governance, security, and accountability. 

Picture a regulated company facing a compliance review six months from now. Who used a given model? What data could it access? Which version generated a specific response? Where was it running, and was that location compliant with data residency requirements? And increasingly, organizations need to know not just where the model is running, but where their data is processed, retained, and governed, particularly as sovereignty requirements differ across countries and industries. What policy was in effect at that moment? Many organizations would struggle to answer all of those consistently today, not because they're careless, but because the tooling to track it hasn't caught up to how fast AI adoption has moved. 

A mature operating model builds a chain of accountability by design, including who made the request, what model processed it, what it was permitted to touch, and what policy governed the interaction. Approving a model isn't governance. Knowing how that model behaves in everyday production is. 

Identity Has to Follow the Request

Nowhere does that accountability matter more than access control, and shadow AI is where the risk shows up first. In fact, “Gartner® predicts that by 2030, more than 40% of enterprises will experience security or compliance incidents linked to unauthorized shadow AI.”[1] 

When employees stand up their own models or reach for consumer-grade frontier tools without oversight, enterprise data can end up flowing to places IT never approved. The model shouldn’t become a way to circumvent the identity and access controls the organization already has in place. 

The fix isn't a new set of AI-specific permissions bolted on after the fact. It's making sure a user's existing identity, role, and access rights travel with every request the model handles on their behalf. If someone can't pull a particular financial record or HR file directly, asking an AI agent the right question shouldn't get them there anyway.  

This gets more urgent as agents move from answering questions to taking action. That's one of the fundamental differences with agentic AI: we're moving from systems that generate answers to systems that can take actions on our behalf. An agent that can modify a ticket, kick off a workflow, or touch a database isn't just retrieving information anymore, it's acting with a level of authority that needs the same guardrails as any employee. The principle is simple to state and easy to violate without the right architecture. AI should not hand someone privileges they didn't already have. 

Models Will Churn, but Your Applications Shouldn't Have To

Even with governance and routing sorted out, there's a structural problem underneath all of it, which is that models are moving faster than almost anything IT has had to absorb before. A model that's the right choice today may not be the right choice in six months, whether because a better one shipped, costs shifted, or a new use case demands different capabilities. 

We've solved similar problems in infrastructure before. Virtualization decoupled applications from the hardware underneath them. Kubernetes did something similar for the infrastructure layer beneath cloud-native apps. Nobody rebuilds their environment every time a new server generation ships, and AI applications shouldn't need a redesign every time a better model arrives. 

An application tightly coupled to one model, API, or deployment environment turns every model change into a migration project. A platform that absorbs that churn instead lets you swap models, infrastructure, or deployment locations with minimal impact to the application built on top. That's the difference between an AI platform that can evolve and one that requires another migration every time the technology changes. 

From Managing Resources to Managing Intent

At enterprise scale, the challenge becomes fragmentation. A study by Gartner® estimates that ”Over 40% of agentic AI projects will be canceled by the end of 2027, due to escalating costs, unclear business value or inadequate risk controls”. [2] Those are also the kinds of challenges a fragmented AI environment can amplify.  

Fragmentation is manageable with two AI projects running somewhere in the organization. It becomes an entirely different problem at a few hundred, spread across public cloud, on-prem GPUs, and the edge, each with its own tooling for serving, security, and monitoring. IT ends up inheriting duplicated teams, inconsistent audit trails, and no clear picture of what's running where or what data it can touch. 

A consistent operating model doesn't mean forcing everything into one location. AI is going to be distributed by nature, across data centers, cloud, and the edge. That's ultimately what an enterprise AI factory needs to provide: a repeatable way to bring together infrastructure, models, data, security, governance, and operations without locking each AI application to a specific technology stack. What a consistent control plane provides is consistency across that diversity, whichever environment a given workload lands in. 

Instead of IT managing every GPU, model, and cluster individually, the platform should understand the outcome you are trying to achieve and handle more of those decisions on its own.  The next stage of enterprise AI isn't about managing more resources more carefully. It's about managing intent and policy, and letting the platform increasingly handle the infrastructure decisions needed to deliver the desired outcome. 

To learn more, visit: https://www.nutanix.com/enterprise-agentic-ai 

To dive deeper, visit Tech Insights on The Forecast 

[1] Gartner press release, Gartner Identifies Critical GenAI Blind Spots That CIOs Must Urgently Address, November 19, 2025, https://www.gartner.com/en/newsroom/press-releases/2025-11-19-gartner-identifies-critical-genai-blind-spots-that-cios-must-urgently-address0 

[2] Gartner press release, Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027, June 25, 2025, https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027 

GARTNER is a registered trademark and service mark of Gartner, Inc. and/or its affiliates in the U.S. and internationally and is used herein with permission. All rights reserved.

©2026 Nutanix, Inc. All rights reserved. Nutanix, the Nutanix logo and all Nutanix product and service names mentioned herein are registered trademarks or trademarks of Nutanix, Inc. in the United States and other countries. Kubernetes is a registered trademark of The Linux Foundation in the United States and other countries. All other brand names mentioned herein are for identification purposes only and may be the trademarks of their respective holder(s). Certain information contained in this publication may relate to, or be based on, studies, publications, surveys and other data obtained from third-party sources. While we believe these third-party studies, publications, surveys and other data are reliable as of the date of this paper, they have not independently verified unless specifically stated, and we make no representation as to the adequacy, fairness, accuracy, or completeness of any information obtained from third-party sources.