The AI Infrastructure Race Is Shifting from GPU Capacity to Operational Execution

July 30, 2026
Angela Shugarts
Angela Shugarts
No items found.

The global buildout of AI infrastructure is accelerating at extraordinary speed. Billions of dollars are flowing into GPUs, data centers, networking, storage, and sovereign computing initiatives. Yet acquiring infrastructure is only the beginning.

The harder challenge is turning that infrastructure into a secure, self-service platform that customers can use and providers can monetize.

In a recent theCUBE Conversation, Rafay Systems Co-founder and CEO Haseeb Budhani joined theCUBE host John Furrier to discuss the rapid expansion of AI infrastructure, the rise of regional and sovereign AI clouds, and why operational execution has become one of the most important differentiators in the market.

Sovereignty Is Reshaping the AI Infrastructure Market

AI infrastructure is no longer being concentrated exclusively within a small number of hyperscale cloud providers. Regional cloud providers, telecommunications companies, NeoClouds, and sovereign AI initiatives are emerging across the world.

Budhani described sovereignty of compute and data as one of the strongest drivers behind this expansion. Organizations increasingly want AI workloads, models, and sensitive data to remain within a specific country or region.

This is also changing how enterprises evaluate infrastructure providers. Many enterprises may continue using public clouds for traditional workloads while choosing local or regional providers for AI initiatives where data control, regulatory requirements, latency, or national interests are more consequential.

That creates a significant opportunity for regional providers, but only when they can meet enterprise expectations.

Organizations expect more than nearby GPU capacity. They require secure access, quotas, auditability, policy controls, workload isolation, and a reliable consumption experience. Building those capabilities takes considerable time and engineering effort.

A GPU Cluster Does Not Automatically Become a Cloud

One of the interview’s strongest themes was the distinction between infrastructure hosting and a true cloud experience.

A cloud must enable a user to access a portal or API, select a service, and receive it on demand without relying on manual intervention. That service must also be delivered securely across multiple customers and teams.

Without self-service and multi-tenancy, providers are often operating custom infrastructure environments rather than scalable cloud platforms.

This distinction matters because manually building a separate environment for every customer is costly and slow. Each deployment creates additional configuration work, operational overhead, and delay between customer demand and revenue recognition.

Hyperscalers spent many years and employed enormous engineering organizations to develop their consumption models. New AI infrastructure providers do not have the luxury of following the same timeline. Their GPUs are already deployed, depreciating, and expected to generate returns.

The competitive question is therefore not whether a provider could eventually build every required capability internally. It is whether the provider can bring a dependable service to market before customers move elsewhere.

Multi-Tenancy Must Extend Across the Entire Stack

Multi-tenancy is often reduced to a Kubernetes configuration problem. In practice, it must be addressed across the full infrastructure stack.

That includes:

  • Bare metal
  • Networking and storage
  • Kubernetes and virtual machines
  • Identity and access management
  • APIs and authentication
  • Usage controls and quotas
  • Observability and audit trails

The complexity grows further when each enterprise customer contains multiple business units, teams, projects, and individual users.

Providers must be able to securely support competing organizations on shared infrastructure while preserving isolation, performance, governance, and cost accountability. These requirements are easy to underestimate during the early stages of an AI cloud build, but they quickly become decisive as the platform begins onboarding enterprise customers.

For Rafay, this is where years of cloud-native operating experience translate directly into the AI infrastructure market. The underlying requirements for safe, governed infrastructure consumption remain relevant, although the financial stakes, workload density, and pace of execution are considerably higher.

Successful AI Clouds Need a Portfolio of Services

The AI infrastructure market is also moving beyond a single consumption model.

Some customers want long-term bare-metal capacity. Model builders may prefer Kubernetes or SLURM clusters. Enterprise teams may expect virtual machines, workbenches, or serverless experiences. Other customers increasingly want to consume open-source models through familiar APIs without managing any infrastructure directly.

A successful AI cloud must support this range of demand.

Budhani characterized this as a portfolio approach. Lower-level infrastructure services can provide stable, long-term revenue, while packaged compute and AI services can deliver greater value and potentially higher margins.

The opportunity expands as providers move from selling:

  1. Raw infrastructure
  2. GPU access
  3. Packaged compute environments
  4. Models, inference, and tokens as services

Each layer requires greater operational sophistication, but it also allows providers to address a broader customer base and capture more value from the same infrastructure investment.

Moving from GPU Hours to Token Consumption

Open-source model delivery is becoming an important use case for both enterprises and AI cloud providers.

Organizations may continue using leading proprietary models while also seeking open-source alternatives for cost control, capacity flexibility, data governance, or workload-specific requirements. They want those models delivered with the same straightforward experience they receive from established AI API providers.

Rafay’s Token Factory is designed to support that model. It enables organizations to deliver open-source models as governed services through APIs, with token-level consumption visibility and controls.

Instead of requiring developers to request clusters or manage model-serving infrastructure, the provider can give them an endpoint and an allocation of tokens. The underlying platform manages infrastructure access, tenant boundaries, governance, and usage.

For an enterprise, this can support internal distribution of AI services across departments and teams. For a NeoCloud or regional provider, it creates an opportunity to sell token-metered services to multiple enterprise customers.

The larger shift is from exposing infrastructure to delivering an immediately consumable AI capability.

Delivery Speed Has Become a Core Product Capability

Technology alone is insufficient in a market moving from monthly implementation cycles to deployment expectations measured in days.

During the conversation, Budhani emphasized that Rafay’s delivery organization has become as strategically important as its software. Customers need help progressing from a signed agreement to a working service, onboarding initial tenants, and responding to real customer requirements.

That demands close coordination among infrastructure vendors, software providers, integrators, and the customer’s own engineering teams.

Rafay increasingly engages customers early in the infrastructure procurement process so the platform and service model can be designed around the intended buyers. The critical questions are commercial as well as technical:

  • Who will consume the infrastructure?
  • Which services will they purchase?
  • What controls will they require?
  • How will they be onboarded?
  • How quickly must the first offering reach the market?

In a rapidly developing market, several months can represent a material competitive disadvantage. Providers must be able to launch an initial service quickly and expand the portfolio as demand becomes clearer.

AI Infrastructure Is an Ecosystem Business

No single company supplies every component required to build and operate a production AI cloud.

A deployment may involve GPU manufacturers, server OEMs, networking companies, storage providers, security vendors, software platforms, global systems integrators, and regional implementation partners.

Supply constraints can also force providers to operate heterogeneous environments. A cloud may use equipment from several server or networking vendors, each with implementation differences that orchestration software must accommodate.

This makes interoperability and technical collaboration essential.

Budhani discussed Rafay’s work with NVIDIA and infrastructure partners including Dell, Cisco, Lenovo, HPE, WWT, storage providers, security companies, and implementation partners. These relationships extend beyond conventional channel agreements. They involve technical alignment, certification, co-design, integration, and coordinated delivery.

The operating platform sits at the intersection of these systems, making partner collaboration central to both the product architecture and the go-to-market model.

Security and Governance Cannot Be Deferred

The pressure to move quickly does not reduce enterprise security expectations.

Distributed AI infrastructure introduces significant requirements around identity, remote access, workload isolation, networking, auditability, and infrastructure control. Providers need security mechanisms that span the data center, cluster, application, and API layers.

These controls must be designed into the platform rather than added after customers begin running sensitive workloads.

Rafay’s approach includes zero-trust access and multi-tenancy across infrastructure, networking, storage, Kubernetes, virtual machines, authentication, tokens, and API gateways. The platform also integrates with ecosystem partners for specialized security requirements outside its direct scope.

For regional and sovereign providers, these capabilities help establish the enterprise trust needed to serve financial services, pharmaceutical, government, telecommunications, and other regulated customers.

From Infrastructure Buildout to Sustainable AI Business

The AI infrastructure market is moving into a new phase.

GPU availability remains important, but capacity alone is becoming less differentiated. Buyers increasingly evaluate how easily they can access services, how securely workloads are isolated, how consumption is measured, and how quickly providers can respond to new requirements.

The providers that succeed will need to combine infrastructure scale with:

  • Cloud-grade self-service
  • Multi-layer multi-tenancy
  • Enterprise governance
  • Flexible service catalogs
  • Usage visibility and metering
  • Operational resilience
  • Rapid customer onboarding
  • A deeply integrated partner ecosystem

As Furrier summarized during the conversation, AI infrastructure must be operated and monetized, not merely installed.

For Rafay, the current market represents the convergence of years spent solving cloud-native operational challenges with an urgent new requirement: helping organizations turn rapidly expanding GPU estates into usable, governed, and revenue-generating AI platforms.

The infrastructure race is accelerating. The next competitive advantage will come from how effectively that infrastructure is operated.

Share this post

Want a deeper dive in the Rafay Platform?

Book time with an expert.

Book a demo

You might be also be interested in...

News

NVIDIA ❤️ Rafay Managed Kubernetes

Rafay MKS has achieved NVIDIA GPU Operator partner validation, providing platform teams with a standardized, governed approach to deploying GPU-accelerated Kubernetes. Learn how to move beyond manual, inconsistent configurations to repeatable, version-controlled AI infrastructure using Rafay Cluster Blueprints.

Read Now

Product

Telco Cloud Services: How Communication Service Providers Can Monetize Cloud Infrastructure

Read Now

Product

One-Click Digital Twins: Deploying the NVIDIA Omniverse DSX Blueprint using Rafay

See how Rafay transforms the NVIDIA Omniverse DSX Blueprint into a one-click, self-service digital twin offering with governed GPU access, multi-tenancy, and automated session management.

Read Now