I spent this week at the AI Infrastructure Summit in Santa Clara, three days, eight stages, 400+ speakers, 250+ partners, and thousands of conversations and one session reframed how I think about the entire AI buildout. It was a joint talk from Vijay Sairam Pratap, who works with the DSX team at NVIDIA, and my colleague Mohan from Rafay. I walked in expecting a speeds-and-feeds update. I walked out convinced that the most important shift happening in this market isn't technical at all, it's financial.
The question that opened the session
Vijay opened with a deceptively simple framing: are we assembling data centers, or are we building something fundamentally different? His argument was that the demands on AI infrastructure have changed dramatically in just the last 36 months, workloads have become multimodal, agentic, and workflow-driven, and that you can no longer simply assemble parts. You have to integrate every layer, from how energy flows in, to the chips that convert that energy into tokens.
That's where the phrase that stuck with me came from: compute is the new revenue. The AI factory is the industrial infrastructure of the next era, a place where power flows into a factory, gets converted into compute, and that compute generates the tokens that power every downstream application.
The part that made the CFOs in the room lean in
The framing that really landed for me came later, from Mohan. His point was blunt: this is no longer an IT decision. When you're spending $100 million, your CFO is going to ask you to run a model showing how and when you make money. It's a financial decision, not a technical one.
He laid out a monetization ladder that I think every operator writing a nine-figure check should have on the wall. Four tiers, each climbing the value chain:
Wholesale land, power, and shell. You wrap the building and sell access to a single large buyer. The margins here are low single digits. A commodity business.
GPU-as-a-Service. Instead of one tenant consuming everything, you diversify -> 20 tenants, 100 tenants, on long-term or shorter monthly commits.
Compute-as-a-Service. Now you start to look like a hyperscaler. Users log in, click a button, and get a GPU virtual machine, a Slurm cluster, or a Kubernetes cluster. They do less work, so they'll pay more for it.
AI token factory. The most sophisticated model. You stop selling infrastructure entirely and sell tokens—billions of them—and if you run it efficiently, margins can be incredibly high.
The keyword, he repeated, was efficiently. The same infrastructure run at 50% utilization versus high efficiency is the difference between a decent return and a dramatically better one.
Why this isn't theoretical
The number that made the room go quiet: Mohan described a customer that started a year ago with 500 GPUs and today runs 150,000. That kind of growth breaks things, observability collapses, and you have no idea what's failing across your data centers. His point was that the operators who win aren't the ones with the most GPUs; they're the ones who can climb the value chain without their operations falling apart.
And the economics are what force the climb. For a customer with 150,000 GPUs, a 5% margin doesn't justify the investment, they need to operate at 20% or 30%.
What I took away
Vijay framed the moment as bigger than a product cycle. Mohan put it personally: he lived through the '99 crash, and this feels different, the market is going to look very different over the next few years.
Here's my own takeaway from sitting there: the GPU rental business still exists, but it's already a commodity. The real question every operator should be asking isn't "how many GPUs can I deploy?"—it's "how far up the value chain can I get, and how efficiently can I run once I'm there?" Compute is the new revenue. The winners will be the ones who treat it that way.
52.6% More Tokens, 1.53× Less GPU Time: Production Inference with Minima and Rafay
A single NVIDIA Blackwell GPU, an optimized Qwen3.6-27B worker from Minima, and Rafay Token Factory around it. The result was a governed, multi-tenant, token-metered service that delivered 52.6% more output than a strong FP8 baseline, with usage records reconciled at 99.97% for billing
The AI Infrastructure Race Is Shifting from GPU Capacity to Operational Execution
Rafay CEO Haseeb Budhani joins theCUBE to discuss sovereign AI, multi-tenant cloud platforms, GPU monetization, and the shift from infrastructure to AI services.