Rafay Platform Achieves NVIDIA-Certified Hypervisors Status
Rafay's VMaaS offering is now NVIDIA-Certified for HGX and NVL72 systems, giving operators secure multi-tenancy at near bare-metal GPU performance.
Read Now

Rafay MKS, the Managed Kubernetes Service within the Rafay Platform, has completed NVIDIA’s GPU Operator partner-validation program.
The validated configuration and deployment procedure are now documented on the NVIDIA partner-validation documentation, giving organizations a repeatable path for deploying GPU-accelerated Kubernetes clusters with Rafay MKS.
The validated configuration can be packaged within Rafay as a version-controlled Add-On, incorporated into a reusable Cluster Blueprint, and applied consistently across GPU-enabled Kubernetes fleets.

That helps organizations move from manually configured GPU clusters to a standardized operating model for AI infrastructure.
The validation was completed using a production-representative Rafay MKS cluster.
The NVIDIA GPU Operator was deployed through Rafay as a Helm-based Add-On using the standard NVIDIA chart and a pinned driver version.
No MKS-specific RBAC changes, elevated privilege modifications, or container-runtime customizations were required. Rafay MKS uses CNCF-conformant upstream Kubernetes, allowing the NVIDIA software stack to be deployed using familiar Kubernetes and Helm conventions.
Installing the NVIDIA GPU Operator manually may be manageable for one cluster. Operating it consistently across dozens or hundreds of clusters presents a different challenge.
Versions change. Configuration values diverge. Drivers drift. Upgrade timing varies. Over time, clusters that were intended to be identical become operationally inconsistent.
Rafay addresses that problem by allowing platform teams to define the NVIDIA GPU Operator once as a versioned Add-On and incorporate it into a Cluster Blueprint.

The Blueprint can specify:
When a GPU-enabled cluster is provisioned, the Rafay control plane reconciles the cluster against its assigned Blueprint and deploys the approved configuration automatically.
Platform teams can therefore treat the validated NVIDIA configuration as a governed platform artifact rather than a sequence of commands that must be repeated manually on every cluster.
When an upgrade is approved, teams update the Add-On centrally and roll out the change through controlled Blueprint reconciliation.
The NVIDIA GPU Operator leverages the Kubernetes Operator Framework to automate the end-to-end management of the NVIDIA software stack. By containerizing critical components, including NVIDIA drivers, the Kubernetes device plugin, the NVIDIA Container Toolkit, and monitoring tools, it simplifies initial deployment and ongoing lifecycle management.
Architecturally, the operator runs as a Custom Resource Definition (CRD) with a controller that manages component state, offering flexible installation options that allow it to run within any namespace while maintaining standard dashboard compatibility.

NVIDIA GPU Operator automates the software components required to expose NVIDIA GPUs to Kubernetes workloads, including drivers, device plugins, container-runtime integration, monitoring components, and related services.
Within the Rafay Platform, that GPU-aware Kubernetes foundation becomes part of a broader operating model for AI infrastructure.
Rafay enables organizations to:
GPU Operator makes Kubernetes aware of the underlying NVIDIA GPUs. Rafay makes those GPU-enabled environments repeatable, governed, and consumable across teams and tenants.
AI infrastructure rarely remains a single-cluster project.
As usage expands, organizations must support multiple teams, workloads, GPU types, software versions, security requirements, and upgrade schedules. Manual configuration creates operational risk precisely when the infrastructure becomes most valuable.
A validated and Blueprint-driven approach provides platform teams with:
Faster deployment
GPU-enabled Kubernetes environments can be provisioned with the required NVIDIA software stack already defined.
Configuration consistency
Operator versions, drivers, namespaces, and Helm values can be standardized across clusters.
Controlled upgrades
Platform teams can test, approve, and roll out NVIDIA software updates through governed platform workflows.
Reduced configuration drift
Clusters are continuously reconciled against the approved Blueprint rather than relying on one-time installation procedures.
A foundation for self-service AI
Validated GPU-enabled clusters can support developer-facing Kubernetes environments, AI workbenches, inference services, and packaged applications without exposing end users to the underlying infrastructure complexity.
Existing Rafay customers can deploy the validated configuration by:

This milestone represents another step in Rafay’s broader work with NVIDIA to help organizations operationalize accelerated infrastructure.
Rafay will continue validating additional configurations as NVIDIA’s hardware and software ecosystem evolves. The teams are also collaborating across a wider set of AI infrastructure initiatives, including NVIDIA AI Enterprise software, NVIDIA NIM microservices, inference services, and AI Factory operating models.
For customers, the objective is straightforward: reduce the engineering effort required to move from installed GPUs to secure, repeatable, production-ready AI environments.
Organizations building GPU infrastructure can use the Rafay Platform to coordinate that journey, from infrastructure and Kubernetes operations through multi-tenant consumption and AI service delivery.
To see the validated configuration or explore how Rafay operationalizes NVIDIA-accelerated infrastructure, contact the Rafay team.
Thank you to the NVIDIA GPU Operator team for their guidance throughout the validation process.

Rafay's VMaaS offering is now NVIDIA-Certified for HGX and NVL72 systems, giving operators secure multi-tenancy at near bare-metal GPU performance.
Read Now
Rafay CEO Haseeb Budhani joins theCUBE to discuss sovereign AI, multi-tenant cloud platforms, GPU monetization, and the shift from infrastructure to AI services.
Read Now

See how Rafay transforms the NVIDIA Omniverse DSX Blueprint into a one-click, self-service digital twin offering with governed GPU access, multi-tenancy, and automated session management.
Read Now