NVIDIA ❤️ Rafay Managed Kubernetes

July 24, 2026

Rafay MKS, the Managed Kubernetes Service within the Rafay Platform, has completed NVIDIA’s GPU Operator partner-validation program.

The validated configuration and deployment procedure are now documented on the NVIDIA partner-validation documentation, giving organizations a repeatable path for deploying GPU-accelerated Kubernetes clusters with Rafay MKS.

The validated configuration can be packaged within Rafay as a version-controlled Add-On, incorporated into a reusable Cluster Blueprint, and applied consistently across GPU-enabled Kubernetes fleets.

That helps organizations move from manually configured GPU clusters to a standardized operating model for AI infrastructure.

What NVIDIA and Rafay validated

The validation was completed using a production-representative Rafay MKS cluster.

The NVIDIA GPU Operator was deployed through Rafay as a Helm-based Add-On using the standard NVIDIA chart and a pinned driver version.

No MKS-specific RBAC changes, elevated privilege modifications, or container-runtime customizations were required. Rafay MKS uses CNCF-conformant upstream Kubernetes, allowing the NVIDIA software stack to be deployed using familiar Kubernetes and Helm conventions.

From a validated cluster to a repeatable fleet standard

Installing the NVIDIA GPU Operator manually may be manageable for one cluster. Operating it consistently across dozens or hundreds of clusters presents a different challenge.

Versions change. Configuration values diverge. Drivers drift. Upgrade timing varies. Over time, clusters that were intended to be identical become operationally inconsistent.

Rafay addresses that problem by allowing platform teams to define the NVIDIA GPU Operator once as a versioned Add-On and incorporate it into a Cluster Blueprint.

The Blueprint can specify:

  • the NVIDIA GPU Operator Helm repository;
  • the approved GPU Operator version;
  • the target namespace;
  • the required driver version; and
  • environment-specific Helm values.

When a GPU-enabled cluster is provisioned, the Rafay control plane reconciles the cluster against its assigned Blueprint and deploys the approved configuration automatically.

Platform teams can therefore treat the validated NVIDIA configuration as a governed platform artifact rather than a sequence of commands that must be repeated manually on every cluster.

When an upgrade is approved, teams update the Add-On centrally and roll out the change through controlled Blueprint reconciliation.

Understanding the NVIDIA GPU Operator

The NVIDIA GPU Operator leverages the Kubernetes Operator Framework to automate the end-to-end management of the NVIDIA software stack. By containerizing critical components, including NVIDIA drivers, the Kubernetes device plugin, the NVIDIA Container Toolkit, and monitoring tools, it simplifies initial deployment and ongoing lifecycle management. 

Architecturally, the operator runs as a Custom Resource Definition (CRD) with a controller that manages component state, offering flexible installation options that allow it to run within any namespace while maintaining standard dashboard compatibility.

The role of GPU Operator in the Rafay Platform

NVIDIA GPU Operator automates the software components required to expose NVIDIA GPUs to Kubernetes workloads, including drivers, device plugins, container-runtime integration, monitoring components, and related services.

Within the Rafay Platform, that GPU-aware Kubernetes foundation becomes part of a broader operating model for AI infrastructure.

Rafay enables organizations to:

  • provision and manage Kubernetes clusters across data centers, cloud, and hybrid environments;
  • deliver Kubernetes, virtual machines, bare metal, SLURM clusters, and AI environments through common platform workflows;
  • apply multi-tenant access controls, quotas, policies, and auditability;
  • publish standardized compute and application offerings through self-service catalogs;
  • track infrastructure consumption by tenant, team, or workload; and
  • deliver higher-level AI services, including notebooks, inference endpoints, NVIDIA NIM microservices, and token-metered model APIs.

GPU Operator makes Kubernetes aware of the underlying NVIDIA GPUs. Rafay makes those GPU-enabled environments repeatable, governed, and consumable across teams and tenants.

Why this matters for AI infrastructure teams

AI infrastructure rarely remains a single-cluster project.

As usage expands, organizations must support multiple teams, workloads, GPU types, software versions, security requirements, and upgrade schedules. Manual configuration creates operational risk precisely when the infrastructure becomes most valuable.

A validated and Blueprint-driven approach provides platform teams with:

Faster deployment
GPU-enabled Kubernetes environments can be provisioned with the required NVIDIA software stack already defined.

Configuration consistency
Operator versions, drivers, namespaces, and Helm values can be standardized across clusters.

Controlled upgrades
Platform teams can test, approve, and roll out NVIDIA software updates through governed platform workflows.

Reduced configuration drift
Clusters are continuously reconciled against the approved Blueprint rather than relying on one-time installation procedures.

A foundation for self-service AI
Validated GPU-enabled clusters can support developer-facing Kubernetes environments, AI workbenches, inference services, and packaged applications without exposing end users to the underlying infrastructure complexity.

Getting started

Existing Rafay customers can deploy the validated configuration by:

  1. Provisioning an MKS cluster with one or more GPU-enabled worker nodes.
  2. Creating a Rafay Add-On that references the NVIDIA GPU Operator Helm repository and validated chart version, deployed to the gpu-operator-resources namespace.
  3. Applying the approved driver and configuration values.
  4. Adding the Add-On to a Cluster Blueprint.
  5. Assigning that Blueprint to the appropriate GPU-enabled clusters.


Extending the validated foundation

This milestone represents another step in Rafay’s broader work with NVIDIA to help organizations operationalize accelerated infrastructure.

Rafay will continue validating additional configurations as NVIDIA’s hardware and software ecosystem evolves. The teams are also collaborating across a wider set of AI infrastructure initiatives, including NVIDIA AI Enterprise software, NVIDIA NIM microservices, inference services, and AI Factory operating models.

For customers, the objective is straightforward: reduce the engineering effort required to move from installed GPUs to secure, repeatable, production-ready AI environments.

Organizations building GPU infrastructure can use the Rafay Platform to coordinate that journey, from infrastructure and Kubernetes operations through multi-tenant consumption and AI service delivery.

To see the validated configuration or explore how Rafay operationalizes NVIDIA-accelerated infrastructure, contact the Rafay team.

Thank you to the NVIDIA GPU Operator team for their guidance throughout the validation process.

Share this post

Want a deeper dive in the Rafay Platform?

Book time with an expert.

Book a demo
Tags:

You might be also be interested in...

Product

Rafay and NVIDIA DSX OS: Turning Open-Source Components into a Consumable AI Cloud

AI factory operators have solved the GPU capacity question. The harder one is turning that capacity into production AI services. Rafay integrates NVIDIA DSX OS to ship it as a consumable AI cloud.

Read Now

Product

Automated GPU Health Monitoring with NVIDIA NVSentinel on the Rafay Platform

Every GPU node monitored. Faulty nodes automatically quarantined and remediated. The Rafay Platform and NVIDIA NVSentinel make that a fleet-wide guarantee, not a per-cluster aspiration.

Read Now

Product

The Telco AI Imperative: From Connectivity to Sovereign AI Infrastructure

The AI buildout demands exactly what telcos already have, now is the moment to make that infrastructure count.

Read Now