Rafay AI Summit Day 1: Our Biggest Takeaways on Building Profitable AI Businesses

Rafay AI Summit Day 1: Our Biggest Takeaways on Building Profitable AI Businesses
Today we kicked off the Rafay AI Infrastructure Leadership Summit in Barcelona, bringing together nearly 150 attendees from 27 countries — including NeoCloud operators, infrastructure providers, technology partners, investors, and AI leaders.
Day 1 focused on a question that is becoming increasingly important across the AI infrastructure industry:
How do you turn rapidly expanding AI capacity into a profitable AI business?
The day started with Rafay CEO Haseeb Budhani setting the stage around the shift from simply acquiring infrastructure to building sustainable businesses around it. From there, the conversation moved across the AI infrastructure stack — from NVIDIA's evolving AI factory architecture and networking automation to GPU utilization, power constraints, multi-tenancy, and the rapid growth of inference demand.

There was plenty to unpack.
Here’s a roundup of our biggest takeaways from Day 1 of the Rafay AI Summit
1. Capacity is not a business
For the past several years, much of the NeoCloud conversation has revolved around one question:
Who can get the GPUs?
That made sense when capacity was extraordinarily scarce.
But Day 1 made it clear that simply owning GPUs is no longer enough to create a differentiated AI business.
As Haseeb Budhani discussed in the opening session, the challenge now is what happens after the infrastructure arrives.
How quickly can you bring it online?
How efficiently can you operate it?
How much of that capacity can you keep productive?
How easily can customers consume it?
And how do you turn all of that infrastructure into something people are willing to pay for?
That shift — from capacity acquisition to business execution — became the backdrop for nearly every conversation throughout the day.
2. The hardest part starts after the GPUs arrive
Getting thousands of GPUs into a data center is a significant accomplishment.
Operating them at scale is a different challenge entirely.
NVIDIA's Warren Barkley put it simply:
“The actually hardest part is operating at scale.”

An AI factory is not just a collection of accelerators.
Operators have to coordinate servers, networking, storage, GPU software, Kubernetes, scheduling, security, observability, tenancy, power, and customer-facing services.
And the physical scale can be enormous.
A gigawatt-scale AI factory can require approximately 320,000 kilometers of fiber and between two and four million connectors.
At that level of complexity, infrastructure cannot depend on manual configuration or one-off engineering.
It needs to become repeatable.
That is also why technologies such as NVIDIA DSX SIM are becoming increasingly important. Simulation allows infrastructure teams to model and validate large-scale environments before the physical systems are fully deployed.
The takeaway: the operating model needs to be designed alongside the AI factory itself.
3. Standardize early or pay for it later
One of the strongest operational lessons from Day 1 was the importance of standardizing the infrastructure stack early.
Alex, CEO of Netris, discussed the pace at which AI infrastructure is now being deployed, with approximately 45 clusters brought online over two years and roughly four additional clusters being deployed each month.

At that speed, every exception becomes a problem.
A custom network configuration may seem manageable on cluster one.
By cluster ten, it becomes operational debt.
The same applies to provisioning, security, observability, storage, tenancy, and workload management.
The discussion around networking was particularly clear: NeoClouds should not wait until they reach scale to introduce automation.
They should build the first cluster with scale already in mind.
That makes it possible to increasingly “rubber stamp” the architecture as new environments come online.
Standardize cluster one so cluster ten doesn't become another custom engineering project.
4. GPU utilization may be the biggest profitability lever
One of the most striking numbers discussed on Day 1 was that productive GPU arithmetic utilization across the industry remains around 40% — even during training workloads.
For infrastructure that is this expensive, that represents an enormous opportunity.
A GPU costs money whether it is busy or idle.

Power still has to be delivered.
Financing costs continue.
Networking and storage infrastructure still operate.
The data center still has to be staffed.
That means a relatively small improvement in utilization can have a significant impact on the economics of a NeoCloud.
But the conversation also made clear that utilization isn't simply a GPU scheduling problem.
It depends on the entire system.
Storage performance matters.
Networking matters.
Power matters.
Workload placement matters.
Multi-tenancy matters.
Orchestration matters.
DDN, for example, shared results showing that KV cache optimization could improve inference throughput by approximately 75% on average, with improvements reaching as high as 3.46x for long-context workloads.
The broader takeaway was straightforward:
One of the best ways to improve AI infrastructure economics may be getting more useful work out of the GPUs you already own.
5. Power is becoming a software problem
If there were two scarce resources that came up repeatedly during Day 1, they were GPUs and electricity.
As Chester Reid summarized it:
“Power and GPUs — they're the two things we need to solve for.”

And the power challenge is only getting bigger.
Future AI infrastructure designs discussed during the event are moving toward approximately 650 kilowatts per rack in the 2028–2029 timeframe, with architectures expected to require 800V DC to support those densities.
But the conversation went beyond data center design.
Power increasingly has to become part of infrastructure orchestration.
One example discussed was MaxLPS, a power-smoothing technology that can distribute workload demand across racks and potentially allow operators to support approximately 30–40% more GPUs within the same power envelope.
That introduces a very different kind of scheduling challenge.
Schedulers historically asked:
Where is compute available?
Tomorrow's AI infrastructure may also need to ask:
Where is power available?
Where is thermal capacity available?
What is happening on the network?
Which workloads have priority?
Where can this job run most efficiently?
The boundaries between facilities infrastructure and software infrastructure are quickly disappearing.
6. Agentic AI changes the demand equation
One of the most interesting conversations of Day 1 was about what happens to AI infrastructure demand as applications become increasingly agentic.
Traditional generative AI is usually based on relatively simple interactions.
A person sends a prompt.
The model generates a response.
An agentic system can operate very differently.
Give an agent a goal and it may break the task into multiple steps, retrieve information, call APIs, interact with tools, coordinate with other agents, evaluate results, and continue working until it achieves the intended outcome.
That can dramatically increase token consumption.
A task that might involve roughly 900 tokens in a traditional interaction could potentially require 10,000, 100,000, or significantly more as the workflow becomes increasingly agentic.
Noam from Lenovo captured the shift well:
“The shift is from token maxing to outcome maximizing.”

In other words, users may care less about how many tokens an AI system consumes and more about whether it completes useful work.
At the same time, the cost of inference continues to fall rapidly.
And cheaper inference does not necessarily mean less infrastructure demand.
As discussed during the session, Jevons Paradox suggests the opposite may occur: as a resource becomes less expensive, new use cases emerge and overall consumption increases.
Cheaper tokens.
More AI applications.
More complex agents.
More compute.
The result could fundamentally change how the industry thinks about future AI capacity requirements.
7. Multi-tenancy is becoming an economic requirement
Another recurring Day 1 theme was multi-tenancy.
It is often treated primarily as a security feature.
For NeoClouds, it is increasingly an economic one.
Even infrastructure dedicated to a single large customer can contain multiple tenants in practice:
Development teams.
Production workloads.
Training environments.
Inference environments.
Different applications.
Different business units.
Maintenance operations.
Those workloads need isolation across more than just compute.
True AI infrastructure multi-tenancy needs to extend across:
- GPUs
- Compute
- Storage
- East-west networking
- North-south networking
- Identity and access
- Quotas
- Observability
- Metering and billing
The better a provider can safely share infrastructure, the greater the opportunity to keep GPUs productive.
And that connects directly back to profitability.
Higher utilization requires better sharing. Better sharing requires stronger multi-tenancy.
8. NeoClouds need to move higher in the value chain
Day 1 also raised a bigger question about the long-term NeoCloud business model.
Selling GPU capacity is an important starting point.
But it becomes increasingly difficult to differentiate if multiple providers are selling the same GPU at roughly the same hourly price.
The opportunity is to move from selling raw capacity toward delivering higher-value AI services.
That could mean offering:
Managed Kubernetes.
Managed Slurm.
Virtual machines.
AI development environments.
Training services.
Inference endpoints.
Model catalogs.
Fine-tuning.
Enterprise governance.
Token-based services.
In other words, the commercial model begins to shift from:
GPU-hours
to:
AI services and outcomes.
But offering those services requires much more than hardware.
Operators need self-service provisioning.
Metering.
Rate cards.
Quotas.
Governance.
Multi-tenancy.
Observability.
Automation.
Lifecycle management.
This is where infrastructure operations and business operations begin to converge.
The infrastructure platform increasingly becomes part of the revenue platform.
9. No company can build the AI factory alone
Perhaps the clearest message from Day 1 was the importance of the ecosystem.
The AI factory is simply becoming too complex for one company to own every layer.
NVIDIA provides foundational compute platforms and software.

Server manufacturers industrialize the hardware.
Networking providers automate the fabric.
Storage companies keep accelerators fed with data.
NeoCloud operators bring capacity into individual markets.
Software providers such as Rafay help connect those layers through orchestration, automation, governance, multi-tenancy, and service delivery.
Warren Barkley offered an interesting perspective on how NVIDIA thinks about this.
He described one of the questions Jensen Huang regularly asks:
“Who in the ecosystem is using your stuff, and how happy are they?”
It is a useful way to think about the AI infrastructure market.
The winners may not be the companies that try to build everything themselves.
They may be the companies that build the strongest ecosystem and make those technologies work together as one operating platform.
Day 1 takeaway: From capacity to profitable AI businesses
The conversations on Day 1 reinforced just how quickly the AI infrastructure market is maturing.
The first challenge was getting GPUs.
Now the questions are changing.
How efficiently are those GPUs being used?
How quickly can new clusters be deployed?
How much infrastructure can be automated?
How securely can capacity be shared?
How do you operate within increasingly tight power constraints?
And how do you package all of that infrastructure into services customers want to consume?
That is the next phase of the NeoCloud opportunity.
The goal isn't simply getting from metal to token.
It's figuring out how to get from metal to token to revenue.
That's a wrap on Day 1!
One day down, and there is already plenty to digest.
Day 1 started with the economics of building an AI business and quickly showed how interconnected those economics are with infrastructure design, utilization, networking, power, software, and ecosystem strategy.
The biggest takeaway?
Building AI capacity is only the beginning. Operating it efficiently — and turning it into services customers value — is where the real business gets built.
And with more conversations still ahead in Barcelona, we're just getting started.











