Elevating infrastructure together

Rafay + Red Hat

Joint Reference Architecture for Sovereign AI Cloud

Rafay and Red Hat are partnering to help telecommunications providers, sovereign cloud operators, NeoClouds and enterprises turn GPU infrastructure into governed, self-service AI clouds that their customers can buy with confidence.

By pairing Red Hat AI, built on OpenShift, with Rafay's orchestration and operations and commercial layer, operators get one consistent way to deliver AI services on-premises, in hosted environments and across distributed sites. It combines the enterprise-grade security and platform trust that regulated customers require with the self-service, multi-tenancy, metering and billing needed to turn capacity into revenue. The joint architecture is published in the Red Hat Architecture Center.

Rafay and Red Hat Reference Architecture
Sovereign AI Cloud as a Service

From AI Infrastructure to Scalable, Monetizable AI Clouds

The reference architecture names the components at each layer and traces a tenant request from SKU selection through to a metered GPU environment. Red Hat owns cluster lifecycle, isolation, and policy enforcement across the fleet. Rafay turns those capabilities into services a customer can buy and turns consumption into an invoice.

Joint Solution Components
What It Means for Customers

A joint blueprint from Red Hat and Rafay for turning distributed GPU infrastructure into a governed, self-service, revenue-generating AI cloud.

Rafay Cluster Controller, a Red Hat Certified Operator

Installs on each OpenShift cluster and opens an outbound-only, zero-trust connection to the Rafay control plane, so sites behind NAT or in third-party facilities need no inbound firewall path.

Operators define service offerings that combine GPU capacity, compute profiles, storage, networking, and policy controls, each with rate cards and tenant entitlements attached.

Tenants request and consume services under the operator's own brand, domains, and UI, with tenant-aware IAM and RBAC federated to enterprise identity providers.

Commercial Workflow Engine

Rafay handles approvals, quota validation, capacity reservation, and metering around each request, then calls Red Hat Advanced Cluster Management workflows to fulfill it.

Adds a metered, multi-tenant billing and API layer on top of the vLLM and llm-d inference capabilities in Red Hat AI, so inference can be sold by the token.

Red Hat and Rafay are both NVIDIA AI Cloud Ready validated partners. The architecture schedules against the NVIDIA GPU Operator, Multi-Instance GPU partitioning, and NVIDIA Dynamo, with Red Hat shipping certified versions of the NVIDIA operators as part of the platform lifecycle.

Why the joint architecture matters

Standardize every new site

Red Hat Advanced Cluster Management and OpenShift bring each environment online the same way. Rafay blueprints and drift detection, paired with Red Hat governance policies, hold that baseline across every cluster in the fleet.

Sell what the platform produces

Rafay turns GPU capacity, isolation tiers, and AI services into catalog items with rate cards. Tenants self-serve, and every unit consumed is recorded against the right rate card for invoicing.

Keep sovereignty at the platform layer

Red Hat components are developed upstream as open source, with signed node images and certified NVIDIA operators, giving operators a platform they can inspect and self-support. Rafay governs tenant access and consumption on top.

Build your sovereign AI cloud on Red Hat OpenShift with Rafay

See the joint solution running against your own environment, from tenant request to metered GPU service.