By Arijit Hawlader, Ashwini Vasanth, Nitin Mehra, Ryan Mathews
At Nutanix, we take immense pride in maintaining an average NPS of 90+ over the past 11 years. Delivering this level of experience requires an unwavering focus on our customers, constant innovation, and a relentless drive to improve. As the latest step in our journey, we are empowering customers with a new channel to seamlessly manage their infrastructure by bridging the gap between complex infrastructure questions and rapid, accurate resolutions. To solve this, we built and deployed NIVA (Nutanix Intelligent Virtual Agent), now live on the Nutanix Support Portal.
NIVA is an AI-powered multilingual conversational assistant designed to understand what our customers are trying to accomplish, reason through it against a deep body of product documentation, knowledge base, and guide to a resolution — whether that's a documented answer, a configuration step, or a detailed troubleshooting plan.
Want to see NIVA in action? Watch the NIVA Support Portal Walkthrough on YouTube to see these features live.
Built on Nutanix Enterprise AI (NAI), NIVA serves as an example for customers who are trying to operationalize AI for enterprise-grade use cases. The blog covers the NIVA design philosophy and the reasons for choosing NAI as the solution for inference and model management.
By listening to our customers, we are constantly looking to improve our people, processes, and products. Support traditionally requires users to navigate scattered documentation, creating friction. NIVA, the Nutanix Intelligent Virtual Agent, addresses this friction by providing rapid, natural-language answers sourced from trusted Nutanix documentation and knowledge bases.
Built for customers, partners, and Nutanix employees, NIVA leverages the Agentic AI stack, Nutanix Enterprise AI (NAI), and Nutanix Kubernetes Platform (NKP). We built NIVA on three pillars:
NIVA handles enterprise GenAI complexities such as GPU, security, and accuracy challenges by leveraging an integrated NAI-based stack. This is designed to support high availability and scalability, optimize costs, and avoid vendor lock-in.
NIVA provides near-instant support across the Nutanix product portfolio. Key capabilities include:
Based on the usage patterns and core pain points we’ve observed from the field implementing conversational AI assistants, we wanted to share our design philosophy and choices. If you are currently on your own AI agent-builder journey, we hope our learnings can help guide your path.
The functional and non-functional demands of RAG (Retrieval-Augmented Generation) and agentic applications require dedicated GPU resources for embedding generation, indexing, inferencing, and orchestration. Furthermore, AI applications introduce unique safety and security governance standards that differ significantly from traditional software.
To handle this, NIVA is built to scale independently of the customer-facing and other backend applications. This isolation is designed so that if system complexity grows, driven by future reasoning and expanded agentic capabilities, our core infrastructure remains performant and unburdened by app constraints.
Our architecture relies on having the flexibility to consume open source, as well as frontier models based on the specific needs for a particular use case. Given the rate of evolution of models in the open source as well as the price fluctuations from frontier model providers, we believe this is a pragmatic approach to shield us from vendor lock-in.
For enterprise customer-facing applications, 24/7 uptime isn't a luxury; it’s a mandate. Our applications are designed to target high availability.
NIVA architecture is engineered for robust stability and inherent fault tolerance. It is designed to help our core AI Infrastructure and applications deliver on our rigorous enterprise Service Level Agreement mandates.
Nutanix Enterprise AI (NAI) is a production-grade product, offering the ability to consume open-source, fine-tuned, or provider models securely with lifecycle management capabilities. NAI simplifies the operational side of model deployment and abstracts the intricate details of model inference, namely:
NIVA on Portal evolved through several stages of maturity. Early prototypes relied on standalone Python processes hosting open-source models directly on GPU-enabled systems. While suitable for experimentation, this approach introduced operational challenges around process management, upgrades, scaling, failover, and resource utilization.
The next phase moved the model serving into Kubernetes-managed workloads, improving resiliency and operational consistency. However, model optimization, GPU allocation, version management, and inference tuning still require significant expertise and operational effort.
NAI provided a production-ready platform that abstracts much of the operational complexity associated with model serving. By leveraging NAI, the NIVA development team gained standardized deployment patterns, lifecycle management, health monitoring, authenticated inference endpoints, and operational guardrails required for a customer-facing service. These built-in capabilities reduced the engineering effort needed to securely expose AI services while providing a consistent framework for access control, auditing, and platform operations across models and environments.
Achieving optimal performance from modern open-source models requires tuning numerous inference parameters. Frameworks such as vLLM expose 90+ configuration options that can significantly impact throughput, latency, GPU utilization, and model stability.
Rather than requiring application teams to continuously evaluate and tune these settings for every model release, NAI provides validated deployment configurations and inference best practices out of the box. This allowed the NIVA engineering team to focus on retrieval, agent workflows, and customer experience rather than infrastructure-level model optimization.
Our experience with the default NAI inference configuration has been positive, providing predictable performance, simple operations, and little time required to onboard and evaluate new models.
It is not possible to iterate over all the options necessary to optimize every model as new models are released. NAI provides this out of the box, allowing fast go-to-market solutions with improved customer experiences.
The development of NIVA was a collaborative effort with the Nutanix NAI team. This partnership proved invaluable by incorporating direct usage feedback from the NIVA team. The Nutanix NAI team incorporated a vLLM Sandbox to enable secure experimentation of the vLLM configurations based on the specific use case needs to address the collaboration needs. This sandboxing capability has since been adopted by multiple external customers and has proved valuable.
The Nutanix NAI Team also built key inference optimizations such as vLLM speculative decoding directly into the platform, promoting strong performance and an excellent experience for our customers. This team benefited from the feedback loop to build conviction on capabilities as well as to harden our product offering.
NIVA represents a meaningful step in how Nutanix delivers support, but it's a starting point, not a finish line. We're continuing to expand the knowledge NIVA can draw on, deepen its agentic capabilities across more complex scenarios, and make every interaction faster and more precise. Your feedback is a direct input into that roadmap, so as you use NIVA, tell us what works and what could be better.
We built NIVA because we believe getting help should be as intelligent and seamless as the infrastructure you run on Nutanix. We can't wait for you to try it.
Ready to get started? Sign in to the Nutanix Support Portal and say hello to NIVA.
©2026 Nutanix, Inc. All rights reserved. Nutanix, the Nutanix logo and all Nutanix product and service names mentioned are registered trademarks or trademarks of Nutanix, Inc. in the United States and other countries. Kubernetes is a registered trademark of The Linux Foundation in the United States and other countries. All other brand names mentioned are for identification purposes only and may be the trademarks of their respective holder(s).
This content may contain express and implied forward-looking statements, including but not limited to statements regarding our plans and expectations relating to new product features and technology under development, the capabilities of such product features and technology, and our plans to release product features and technology. Such statements are not historical facts and are instead based on our current expectations, estimates and beliefs. The accuracy of such statements involves risks and uncertainties and depends upon future events, including those that may be beyond our control, and actual results may differ materially and adversely from those anticipated or implied by such statements, including, among others: failure to develop, or unexpected difficulties, delays or disruptions in developing, releasing or distributing, new products, services, product features or technology in a timely or cost-effective basis. Any forward-looking statements included speak only as of the date hereof and, except as required by law, we assume no obligation to update or otherwise revise any such forward-looking statements to reflect subsequent events or circumstances.