Rafay AI Summit Day 2: Our Biggest Takeaways on Inference, Enterprise AI and What Comes Next

September 10, 2026

After a packed first day at the Rafay AI Infrastructure Leadership Summit in Barcelona, we came back for Day 2 with an even bigger question:

What happens when AI infrastructure moves from simply providing GPU capacity to delivering AI services at scale?

Day 2 brought together perspectives from NVIDIA, Accenture, AI infrastructure operators, sovereign AI providers, and Rafay's own product leadership. The conversations moved from optimizing inference and maximizing tokens per watt to enterprise AI adoption, distributed compute, sovereign AI, rack-scale infrastructure, observability, and the operational capabilities NeoClouds will need as they grow.

If Day 1 was about building a profitable AI infrastructure business, Day 2 was about where the next wave of demand is coming from — and how providers can prepare for it.

Here’s a roundup of our biggest takeaways from Day 2 of the Rafay AI Summit

1. The AI factory is increasingly measured in tokens, not GPUs

We spent a lot of Day 1 talking about GPU utilization. On Day 2, NVIDIA took that conversation one step further.

Laura Morselli, Senior Solution Architect at NVIDIA, described the modern AI factory in simple terms: infrastructure that turns power into tokens.

As inference becomes a larger share of AI workloads, the objective is not simply to install more GPUs. It is to produce as many useful tokens as possible from the power, infrastructure, and capital already available.

And that requires looking across the entire stack.

Accelerator efficiency matters. Model architecture matters. Networking matters. Memory matters. Serving software matters. And orchestration matters.

Optimizing only one layer leaves performance — and ultimately revenue — on the table.

The economics become particularly interesting with reasoning and agentic AI.

A user may still see a single answer on the screen, but a reasoning model can generate 10 to 50 times more tokens behind the scenes. An agentic application can go much further, repeatedly planning, calling tools, checking results, and trying again — potentially consuming hundreds of thousands of tokens in a single interaction.

The implication for AI infrastructure providers is significant:

The important metric is becoming less about how many GPUs you operate and more about how much valuable AI work you can produce from them.

2. Software can change the economics of the hardware you already own

One of the more striking examples of Day 2 came from NVIDIA's discussion of Dynamo, its distributed inference serving software.

Inference is not a single uniform workload.

The initial context or “prefill” stage is compute-intensive, while token generation or “decode” is more constrained by memory. Treating those stages differently allows infrastructure to be optimized for each workload instead of forcing both onto the same architecture.

Dynamo separates those phases across GPU resources and combines that with capabilities such as KV-aware routing, scheduling, cache management, and Kubernetes-native orchestration.

The result?

NVIDIA showed an example where software improvements delivered roughly a 4x increase in token throughput on the same hardware compared with an earlier submission.

For a data center that is constrained by power, that type of improvement is economically significant. It is effectively a way to create additional capacity without adding the equivalent amount of physical infrastructure.

This reinforced something we heard throughout the summit:

AI infrastructure optimization is becoming a software problem as much as a hardware problem.

3. Moving up the stack is a business model change — not just a product upgrade

David Wood from Accenture brought an enterprise and business-model perspective to the NeoCloud discussion.

He described the market as moving through three broad phases.

First came the GPU supply problem: How do we get enough compute?

Then came power and capital constraints: How do we finance and power all of this infrastructure?

Now comes the next challenge: How do we capture demand?

The NeoClouds that win may increasingly be those that can shift from thinking primarily as supply-side infrastructure operators to understanding where enterprise AI demand is going and building services around it.

Managed inference and token factories were described as some of the strongest demand signals Accenture is currently seeing.

But moving from bare-metal GPU services into managed inference, platforms, and enterprise AI isn't simply a matter of adding another product SKU.

It changes the business itself.

The sales motion changes.

Support changes.

SLAs change.

Security requirements change.

FinOps changes.

And the organization has to learn how to speak the language of enterprise customers rather than infrastructure buyers.

David summed up the opportunity through an example of a NeoCloud model where 65% of capacity was committed to traditional long-term demand while 35% was reserved for a token factory offering. In that specific model, the token factory layer was projected to generate roughly 2–3x the revenue and 4–5x the margin.

The lesson wasn't that every NeoCloud should follow exactly the same formula.

It was that moving higher in the stack can fundamentally change the economics of the same underlying GPU infrastructure.

4. Enterprise customers don’t want GPUs. They want outcomes.

Enterprise AI adoption generated a lot of discussion on Day 2.

On the surface, there appears to be a contradiction.

Enterprises are consuming huge amounts of AI, yet many organizations still say they haven't realized transformational value from it.

David shared that Accenture itself processes roughly 12.7 trillion tokens per week. But much of today's enterprise usage still centers on areas such as knowledge work, chat, and coding.

The larger transformation may come as agentic AI begins reshaping front- and back-office processes.

Instead of humans being the primary consumers of tokens, millions of applications and agents could generate tokens continuously as they reason, call tools, and perform work.

From that perspective, David described enterprise AI as still being at essentially “inning zero.”

But capturing that opportunity requires NeoClouds to change how they approach enterprise customers.

“You're not selling infrastructure, you're understanding the workload and how that translates into infrastructure.”

That means making AI infrastructure easy to buy, easy to build on, and easy to run.

Enterprises expect provisioning, compliance controls, governance, budgeting, chargebacks, quotas, and cost visibility to already be part of the experience. A large enterprise may technically be one customer while operating thousands of internal tenants, teams, and workloads underneath it.

That is a very different proposition from selling GPU-hours.

5. Sovereign AI has an adoption problem as well as an infrastructure problem

Day 2 also brought the sovereign AI discussion down from policy to practical deployment.

Aras Integrasi shared its experience working with the Malaysian public sector, where the government has articulated ambitions around becoming an AI nation and investing in sovereign AI infrastructure.

The motivation is familiar: organizations want access to modern AI capabilities, but they are also asking fundamental questions about where their data goes and who ultimately controls the underlying infrastructure.

But the session highlighted another side of sovereign AI that gets less attention.

Building sovereign infrastructure does not automatically create sovereign AI adoption.

Government agencies still need applications people actually use.

Legacy systems still need to be integrated.

Employees need training.

Organizations need internal champions.

And leadership needs evidence that AI is improving productivity enough to justify continued investment.

Then there is utilization.

If sovereign GPU infrastructure is restricted to individual agencies or workloads and remains idle much of the week, the economics become difficult.

The speaker highlighted the tension directly: traditional government ownership models may leave expensive infrastructure operating only during normal working hours, creating utilization rates that are difficult to justify. In the example discussed, weekly utilization could fall to around 25%.

That creates an important question for sovereign AI programs:

How do you preserve sovereignty and isolation while still sharing infrastructure efficiently enough to make the economics work?

There is no simple answer yet — but multi-tenancy, governance, and orchestration are clearly going to be part of it.

6. Rack-scale infrastructure is becoming the new unit of operation

The final session of Day 2 brought the conversation back to what Rafay is hearing directly from customers and partners.

One of the clearest trends: rack-scale deployments are becoming much more common.

Instead of treating an environment primarily as individual servers and GPUs, operators increasingly need to treat the entire rack as an operational object.

That changes what infrastructure management needs to understand.

Provisioning.

Networking.

Power.

Cooling.

Health.

Even things like detecting and responding to liquid-cooling leaks.

Rack-scale infrastructure can accelerate deployment and reduce some of the complexity involved in assembling systems component by component, but it also introduces a new level of operational coordination.

As Rafay CPO Mohan Atreya put it toward the end of the session:

“Rack scale is where the action appears to be.”

And with some customers moving from hundreds or thousands of GPUs toward much larger environments, managing the rack as a system is quickly becoming essential.

7. NeoCloud customers expect a hyperscaler-grade experience

Another message Rafay's product team said it is hearing repeatedly from customers is that the user experience bar has already been set by the hyperscalers.

Developers have spent the last 20 years learning what cloud infrastructure should feel like.

They expect APIs.

They expect self-service.

They expect familiar constructs.

They expect to click a button and get a service.

If using a NeoCloud means learning an entirely new way of consuming infrastructure, that creates friction.

That is why NeoCloud differentiation cannot come at the expense of usability.

Rafay shared work underway around more cloud-like APIs, an expanding service catalog, and enterprise controls such as quotas, reservations, chargebacks, governance, and team-level visibility.

The objective isn't to recreate every service offered by a hyperscaler.

It is to give customers a familiar cloud-like experience while exposing infrastructure purpose-built for AI.

That combination could become increasingly important as enterprises use NeoClouds alongside — rather than necessarily instead of — their existing hyperscaler environments.

8. Day 2 operations need to start on Day 1

Perhaps the most practical advice of the entire day came during Rafay's closing session.

Don't wait until you are operating thousands of GPUs to start thinking about observability and automation.

Customers can grow incredibly quickly.

An environment may start with dozens or hundreds of GPUs and expand into thousands, bringing with it servers, storage systems, network controllers, firewalls, cooling systems, and many other moving parts.

Something will eventually fail.

And finding the cause becomes much harder as infrastructure grows.

Rafay shared its approach to going beyond traditional dashboards toward proactive synthetic monitoring across compute, networking, storage, and the infrastructure stack, followed by AI-assisted triage and human-in-the-loop remediation.

The goal is straightforward:

Infrastructure should be able to scale without requiring operations headcount to grow at the same rate.

The team also discussed turning operational knowledge into repeatable playbooks that agents can eventually help execute.

Instead of troubleshooting knowledge sitting in one engineer's head, proven operational procedures can become workflows. Agents can help diagnose an issue, recommend or execute the appropriate playbook, and escalate to a human when the action requires approval.

That feels particularly fitting for an AI infrastructure platform:

use AI not only as the workload running on the infrastructure, but also as a tool for operating the infrastructure itself.

Day 2 takeaway: The next NeoCloud battle is about demand

Across two days of conversations, a clear progression emerged.

The first phase of the AI infrastructure boom was about access to GPUs.

The next challenge became power, capital, and deployment at scale.

Now another phase is emerging:

How do you create more demand, deliver more valuable services, and generate more revenue from the infrastructure you already operate?

That means producing more tokens from every watt.

Using software to increase infrastructure yield.

Building enterprise-ready AI services.

Supporting open models and inference.

Delivering hyperscaler-like experiences.

Finding new sources of power.

Operating sovereign infrastructure efficiently.

And automating operations before scale becomes overwhelming.

The NeoCloud opportunity is expanding well beyond GPU infrastructure.

It is becoming a platform business.

That’s a wrap on Day 2 — and Barcelona!

And with that, we wrapped an incredible couple of days in Barcelona.

What started as an opportunity to bring together customers, partners, NeoCloud operators, infrastructure companies, and AI leaders turned into hundreds of conversations about where this market is going next.

By the end of the event, nearly 400 meetings had been booked among attendees — a good indication of just how much collaboration is happening across this ecosystem.

The technology is moving fast.

The infrastructure is getting bigger.

The economics are changing.

And enterprise demand is only beginning to take shape.

But if there was one theme that connected Day 1 and Day 2, it was this:

The opportunity isn't simply to build more AI infrastructure. It's to make that infrastructure easier to operate, easier to consume, and far more valuable to the customers using it.

That’s a wrap from Barcelona — and just the beginning of what comes next.

Share this post

Want a deeper dive in the Rafay Platform?

Book time with an expert.

Book a demo
Tags:
No items found.

You might be also be interested in...

No items found.