Bringing Intelligent AI Workload Orchestration to Nutanix Kubernetes Platform and Nutanix Enterprise AI

By Debo Dutta, Chief AI Officer, Nutanix and Yiannis Georgiou, Principal Engineer, Nutanix

Today, Nutanix announced its acquisition of Ryax Technologies to further accelerate our vision of empowering enterprises to build, run, and govern agentic AI anywhere. As organizations expand AI workloads across hybrid, multi-cloud, and high-performance computing (HPC) environments, maximizing hardware efficiency and managing compute costs remain top operational priorities. With Ryax’s advanced orchestration, smart scheduling, and telemetry-driven resource optimization targeted for future releases of Nutanix Kubernetes Platform (NKP) and Nutanix Enterprise AI (NAI), Nutanix aims to complement its core platform with even deeper GPU utilization and automated workload placement. Below, we take a closer look at the core infrastructure challenges facing modern enterprise AI and how these planned capabilities are designed to optimize density, dynamic sizing, and cross-cloud execution.

Customer Challenges

Modern enterprise AI spans heterogeneous hardware—CPUs, GPUs, across on-premises and public and neo clouds, on Kubernetes and HPC environments—yet most compute remains severely underutilized. Scaling production AI agents requires clearing three core hurdles:

  • The Cost Wall: Static allocations leave expensive accelerators idle between jobs.
  • Data Gravity: Running models far from enterprise datasets degrades latency and inflates egress costs.
  • Multi-Tenant Governance: Strict isolation requirements conflict with dense resource bin-packing across teams.

All three share a root cause: Infrastructure sizing, cluster selection and resource reservations are often made once, often by hand, and never revisited.

Key Value Proposition

The goal of integrating Ryax into Nutanix Kubernetes Platform (NKP) and Nutanix Enterprise AI (NAI) is to reduce infrastructure fragmentation so developers are able to focus strictly on building. The system would be able to provide:

  • Automated Placement: Intelligently route jobs across multi-cluster, public cloud, and HPC environments.
  • Dynamic Sizing: Replace static allocations with continuous right-sizing driven by historical telemetry.
  • Unified Control: Establish a single native layer to deploy, operate, and scale production AI stacks with minimal friction.
infrastructure diagram

Resource Optimization & GPU Utilization

Nutanix plans to incorporate Ryax capabilities into future releases of NKP, including a resource optimization layer designed to replace static allocations with telemetry-driven, per-execution sizing. These Ryax capabilities include: 

  • Right-Sizing & OOM Auto-Recovery: Learns historical resource profiles to set more precise GPU allocations, automatically catching out-of-memory errors, scaling VRAM, and retrying. Why it matters: This eliminates request padding and job failures without forcing data scientists to act as capacity planners.
  • Fine-Grained GPU Bin-Packing: Provisions fractional GPU slices to pack multiple containerized workloads onto shared hardware. Why it matters: This maximizes GPU access across engineering teams and reduces shadow-IT cloud rentals.
  • Serverless GPUs: Allocates GPU capacity strictly during active compute cycles and releases it immediately upon completion. Why it matters: Eliminates idle time waste, freeing overnight interactive notebook allocations for active training runs.

The impact is real. In testing conducted by Ryax, on a 30-run deep-learning burst, the per-execution sizing reduced node-hours by 62% while finishing 5.7% faster. In a document-intelligence pipeline, serverless allocation reduced GPU hold times from hours to minutes per run, while NVIDIA MIG based fractional GPUs allowed four concurrent executions on a single H100;cutting cost per execution by 52%.

Cost-, Energy-, and Cloud-Aware Smart Scheduling

Ryax will complement NAI with a global meta-scheduling layer to optimize job placement across hybrid, multi-cloud, and non-Kubernetes infrastructure:

  • Intelligent Job Placement: Evaluate training, batch, and inference jobs against performance, cost, and energy goals to select the optimal target node pool. Why it matters: Consolidates pipeline scheduling into a single control plane rather than maintaining isolated stacks.
  • Cross-Cloud & Cross-Platform Optimization: Dynamically route workloads across NVIDIA/AMD fleets, public clouds, and Slurm HPC clusters. Why it matters: Fully utilizing existing on-premises bare-metal and HPC investments before incurring additional cloud spend can help organizations maximize ROI on expensive hardware and minimize cloud opex.
  • Energy-Aware Placement: Score clusters using hardware power models to steer workloads toward the lowest predicted energy footprint. Why it matters: Converts energy management into an automated, reportable operational lever that can support customers’ sustainability objectives.
  • History-Based Autoscaling: Provision precise CPU, memory, and fractional GPU resources based on prior run telemetry. Why it matters: Increases cluster density without requiring developers to rewrite manifests or alter submission workflows.

Combining Nutanix NKP and NAI with Ryax will transform fragmented hybrid infrastructure into a self-optimizing AI engine. Developers and IT Departments want  instant, unified access to models while continuously maximizing GPU utilization, curbing cloud spend, and reducing idle waste behind the scenes. By joining together, Nutanix and its new team members from Ryax aim to meet these needs and more.