Introduction

The landscape of hybrid cloud infrastructure for Nutanix customers is undergoing a fundamental shift. The strategic move from VMware Cloud (VMC) on AWS to Nutanix Cloud Clusters (NC2) on AWS is driven by a clear architectural directive: the requirement for flexibility, structural cost control, and operational sovereignty.

NC2 provides a hybrid multi-cloud platform that provides operational consistency across global AWS regions. By maintaining the same Nutanix Cloud Infrastructure (NCI) software stack and AHV hypervisor on-premises and in the cloud, customers can eliminate the "learning tax" on operations teams and establish a unified fabric for workload mobility and disaster recovery.

In this article we will explore a Nutanix Disaster Recovery design created to provide maximum availability and recoverability across cloud Availability Zones as well as Regions. Within a given AWS region workloads are protected across Availability Zones by AHV Metro Availability and, across AWS Regions, workloads are protected asynchronous by Nutanix Disaster Recovery, allowing the customer to recover their critical workloads irrespective of where the failure occurs.

Primary requirements for the customer design are:

  • Workloads must be recoverable in the following scenarios:
    • Simultaneous failure of 1 or 2 NCI cluster nodes
    • Simultaneous failure of 1 or 2 storage disks
    • AWS Availability Zone failure
    • AWS Region failure
  • Workloads must support being selectively failed over across Availability Zones or Regions
  • Workloads must support test failovers to a DR “bubble network” environment.

Architecture

The environment is built using a "block-and-pod" strategy to ensure predictable scaling and management isolation. A Pod is the fundamental unit of scale, consisting of one to four workload clusters paired with a dedicated management cluster, running the Nutanix NCP management plane, a scaled-out, X-Large instance of Prism Central.

Fig 01. Pod Architecture Fig 01. Pod Architecture

To optimize compliance and resource alignment, the "Block" strategy segregates clusters by licensing use-cases. Dedicated clusters for RHEL, Database (SQL/Oracle), and Windows workloads are included so that core-based licensing is applied efficiently, adhering to ISV vendor licensing requirements. For production environments, the i7ie.metal-48xl bare-metal instance type has been selected as the primary node type to satisfy the high RAM and NVMe storage performance requirements of enterprise-scale workloads.

Resiliency begins within the cluster.  Each cluster is deployed with a minimum of 5 nodes, allowing them to accommodate 2-node/2-disk resiliency, meaning the cluster can lose two nodes simultaneously, or two disks simultaneously, and continue to function without data loss.

This resiliency solution leverages multiple aspects of the Nutanix Cloud Platform to accomplish the customer’s goals:

  • NC2 on AWS
  • AHV Metro Availability
  • Disaster Recovery Asynchronous Replication
  • Nutanix Cloud Manager (NCM)  Self-Service
  • Nutanix Witness Service
  • Nutanix Intelligent Operations
  • AWS VPC and Transit Gateway

Pod Sizing and Scalability Maximums

The architecture adheres to strict scalability limits derived from the Nutanix configuration maximums and customer maintenance window constraints.

Resource Type Configuration Maximum
VMs per Pod 3,500
Nodes per Cluster 21
VMs per Cluster 1,500 (NGT-enabled)
Nodes per Management Subnet 105
Layer-2 Extended Subnets per Pod 100
VTEP Gateway Appliances per Pod 10
Fig 02. Pod Logical Diagram Fig 02. Pod Logical Diagram

AHV Metro Availability (RPO 0)

For mission-critical production workloads, the architecture utilizes AHV Metro Availability to provide automated site-level failover with zero data loss. This provides synchronous replication of workloads from the primary hosting clusters (Pod A) to the secondary hosting clusters (Pod B) across AWS Availability Zones.

Any workload determined to be in-scope for AHV Metro is assigned to a category specifically intended to add it to a corresponding Protection Policy, enabling the VM to begin synchronously replicating to the paired AZ.  Applications which provide native resiliency, such as Microsoft SQL Availability Groups, can be excluded from this configuration.

To prevent split-brain scenarios during an AZ failure, and to automatically initiate Metro availability failover, the Nutanix Witness Service is deployed in a third, geographically remote region (Pod C) to maintain quorum.   The Nutanix Witness Service monitors the Recovery Plan configured for AHV Metro availability and will automatically trigger a failover across Availability Zones if the primary site becomes unavailable. 

If an AHV Metro failover occurs, the network routes for VM subnets must be updated on the AWS Transit Gateway (TG).  This reconfiguration is performed automatically via an NCM playbook which detects the system alert indicating that a witness-initiated failover has occurred and triggers an NCM runbook script to perform the required network routing updates on the TG.  No manual intervention by the customer’s cloud or networking teams is required to enable network connectivity following a failover.

Fig 03. AHV Metro Logical Diagram Fig 03. AHV Metro Logical Diagram

Remote Asynchronous DR

Regional resilience is provided by the remote Disaster Recovery Pod (Pod C), located in a separate geographic region. This tier protects against full regional outages using Asynchronous or NearSynchronous replication configured via Prism Central Protection Policies.

Workload recovery is prioritized according to the customer RPO/RTO tiering model:

Tier RPO RTO
Tier 0 0 - 4 Hours 15 - 60 Minutes
Tier 1 0 - 4 Hours 0 - 4 Hours
Tier 2 0 - 4 Hours 4 - 8 Hours
Tier 3 0 - 24 Hours 8 - 24 Hours
Tier 4 On-Demand 72 Hours+

If the entire production AWS region has become unavailable, customer administrators can initiate a failover of the Recovery Plan on the remote DR pod to bring up the protected workloads on the DR pod.  The Recovery Plan is configured with startup stages so that the workload can be started in the required sequence to ensure dependencies are available as workloads come online.  Network IP addresses of VMs will change upon startup following a failover to addresses that fall within the assigned DR subnets for each region.

Fig 04. Remote DR Logical Design Fig 04. Remote DR Logical Design

Selective Failovers

An additional requirement within this deployment was to be able to selectively fail over a single application to either the AHV Metro secondary AZ (Pod B) or to the Remote DR AZ (Pod C) and have the application be able to run in a production manner in that location.  This requires:

  • Stretched networking across the VM subnets in Pod A and Pod B, to ensure that VMs can run on the same subnet in either location.  This is facilitated by Nutanix Flow Virtual Networking’s subnet extension feature
  • Additional Recovery Plans dedicated to only the workloads contained within the application to be selectively failed over

Validating Architecture

Bubble Test

All three pods within the environment are configured with a second set of Flow Networking VPCs that are designed to mirror the address space of their respective production VPCs, with the key difference that they have no north/south networking configured.  This allows for a safe space, or “bubble”, to which we can perform a Test Failover of the Nutanix Recovery Plans.  This will initiate the creation of all VMs within the Recovery Plan and assign them to the networks within the DR bubble networks.  VM startup, local credential access, and inter- and intra- subnet connectivity can all be validated within the bubble test, all without having to take any production workloads offline.  Once testing is complete, the test VMs are cleaned up and removed by the Recovery Plan’s test cleanup functionality.

Planned Failover

Testing a witness-initiated AHV Metro failover is non-trivial task.  In on-premises environments a common tactic is to simply power off the nodes in the primary cluster and allow the failover to happen.  In an NC2 on AWS environment, this method is not feasible and will result in data loss.  It is not supported to stop or terminate bare metal instances in an NC2 on AWS cluster.

Instead, we use local network firewalls on Prism Element and Prism Central, or in an AWS security group, to block all network communication between the primary pod and both the secondary pod and the witness service.  Once this network connectivity block is in place, a witness-initiated failover will trigger.

Failovers to the remote DR region, as well as selective failovers of applications, can be initiated by the administrators using the Failover mechanism in the appropriate Recovery Plan.

Network Infrastructure

The NC2 architecture leverages a software-defined Overlay (Flow Virtual Networking) sitting atop the AWS infrastructure (Underlay).

The No-NAT Routing Model

As all workloads within the pods need to be accessible to the customer environment at large, the design utilizes a No-NAT routing model. Dedicated network ranges for workloads are assigned to each pod, and routing is configured both within the pod and outside in the customer networking environment to ensure that routing for the application subnets points to the correct destinations.

Fig 05. Network Topology Overview Fig 05. Network Topology Overview

Networking Requirements for High Availability

Maintaining synchronous state across AZs requires a high-performance network underlay.  Network latency between endpoints for synchronous replication must be 5ms or lower.

Component Role in HA Architecture Justification
VPC Peering Direct Synchronous Replication Provides low-latency, point-to-point connectivity between VPCs, bypassing Transit Gateway costs and firewall bottlenecks.
AWS Transit Gateway Managed Failover Routing Enables Prism Central runbooks to automate static route updates and redirect traffic to the active VPC during a failover event.

The VPC-peering link allows the pods to communicate directly with each other without traversing any additional firewalls or routing infrastructure, avoiding latency introduced by additional hops, as well as cost concerns by not traversing metered AWS resources.

The North/South Virtual Machine traffic is managed by an AWS Transit Gateway deployed in the same AWS account as the primary and secondary pods.  Additional permissions are applied to this Transit Gateway (TG) to ensure that the NC2 management security principal can manage the TG’s route table. This becomes a critical task during a AHV metro failover.

Conclusion/Business Value

This architecture demonstrates how Nutanix Cloud Clusters (NC2) on AWS can deliver a highly resilient, enterprise-class disaster recovery solution that protects applications from infrastructure failures ranging from individual node and disk outages to complete AWS Availability Zone and Region failures. By combining AHV Metro Availability for zero-RPO protection within a region with Nutanix Disaster Recovery asynchronous replication for cross-region resilience, organizations can align recovery objectives to application-specific business requirements while maintaining operational simplicity.

From a business perspective, the design reduces operational risk by ensuring critical services remain available during disruptive events and by providing predictable recovery outcomes through automated failover orchestration and recovery plans. Automated networking updates, witness-driven failover decisions, and policy-based protection eliminate many of the manual processes that traditionally increase recovery times and introduce human error during an outage.

The solution also reduces operational complexity by preserving a consistent Nutanix operating model across on-premises and cloud environments. This approach decreases administrative overhead, reduces training and support costs, and streamlines operational processes, enabling IT teams to focus on delivering business value rather than managing disparate infrastructure platforms. 

Finally, the inclusion of bubble-network testing and selective failover capabilities allows organizations to regularly validate disaster recovery readiness without impacting production applications. This improves compliance, increases confidence in recovery procedures, and helps ensure business continuity objectives can be met when they are needed most. The result is a resilient, scalable, and operationally efficient hybrid-cloud platform that helps organizations minimize downtime, reduce risk, control operational costs, and confidently modernize their infrastructure while maintaining the availability, recoverability, and governance required for mission-critical workloads.

 

©2026 Nutanix, Inc. All rights reserved. Nutanix, the Nutanix logo and all Nutanix product and service names mentioned are registered trademarks or trademarks of Nutanix, Inc. in the United States and other countries. All other brand names mentioned are for identification purposes only and may be the trademarks of their respective holder(s). This content reflects an experiment in a test environment. Results, benefits, savings, or other outcomes described depend on a variety of factors including use case, individual requirements, and operating environments, and this publication should not be construed as a promise or obligation to deliver specific outcomes.