A big change is underway in who runs AI. For most of the cloud era, inference lived inside a handful of hyperscalers and frontier model providers. That is no longer strictly true. Telcos, banks, hospital groups, government ministries, and regional hosters are standing up their own AI infrastructure to keep data, cost, and control inside their own walls. TELUS is running a sovereign AI studio in Canada, and similar build-outs are underway across India, Africa, Australia, and Korea.
Every one of these organizations is beginning to operate cloud-like AI infrastructure.. And sooner or later, they all have to answer the same question: how do we charge for this? Most are about to inherit the wrong answer.
One meter increasingly no longer fits
Nearly everything the industry knows about charging for compute traces back to AWS EC2: one line item, billed on machine time. It fit a general-purpose box running a web server, and it survived two decades because of that fit.
It fits a GPU poorly. The difference between a GPU running at 90 percent utilization and one running at 30 percent is most of what an operator is actually selling, and wall-clock time cannot see it. A tenant holding a card at 30 percent pays the same as a tenant driving it at 90 percent, and the operator absorbs the gap. Worse, an operator with one meter can only compete on the one variable it has exposed, which is price per hour. On an infrastructure base where financing comes due whether the cluster runs at 95 percent or 55 percent, that race has no durable winner.
What AI actually lets you charge for
The platform operating the fleet already captures a vast amount of measurable, billable activity produced by AI infrastructure, far surpassing what a simple hourly rate can reflect. Yet, this data rarely finds its way onto an invoice.
Consider a hospital executing medical image inference: its priority isn't the number of GPU minutes expended, but rather ensuring 20,000 scans undergo processing within the established SLA. Similarly, a government agency prioritizes the review of permits, while a telco focuses on latency assurances. Despite relying on the same underlying GPUs, these represent entirely distinct commercial offerings.
As the sector transitions from pricing infrastructure to pricing outcomes, multiple distinct monetization layers are coming to the forefront: outputs, service guarantees, utilization, memory, and energy.
- Compute. Beyond wall-clock time, fully utilized GPU time charges on what the card actually did rather than how long it was held. The efficient tenant is rewarded, and the operator resells reclaimed headroom instead of giving it away. Priority and preemption tiers belong here too.
- Memory and movement. Data moved in and out of GPU memory drives real time and energy on long-context and multimodal work, yet it is invisible on every bill today. Keeping a model warm has a cost, and a warm start is a feature customers will pay for.
- Output. Tokens in and out, priced per model, is what the hosted clouds already use. For image, video, audio, and long documents, response size is the more honest unit. For anything delivered as a finished service, the completed request is cleaner still.
- Power. Energy drawn in kWh is how the layer underneath already works. The AI cloud today absorbs that variability and hands the customer a flat rate, carrying a risk it is not paid for.
- Guarantees. Where inference physically ran, attested. Air-gapped operation. A per-request audit trail. Which models a tenant may reach at all. A jurisdiction guarantee carries a number a bank will pay that a commercial endpoint abroad cannot offer at any price.
The change is that raw infrastructure can sit close to cost while the services on top price on what they are worth. If hours are the only thing being measured, hours are the only thing that can be priced.
Rafay and Solvimon: metering meets billing
One emerging approach is to separate metering from billing. Platforms like Rafay collect trustworthy usage data from AI infrastructure, while specialized billing systems such as Solvimon translate that into commercial models operators can evolve over time.
Rafay is the orchestration and monetization control plane operators run to turn GPU capacity into consumable, governed, billable services. Because Rafay is already scheduling GPUs, enforcing multi-tenancy, and applying policy across bare-metal, virtualized, and Kubernetes environments, it is already observing the signals that make these meters possible. It meters usage at the tenant, team, and workload level and keeps entitlements separate from provisioning. What it does not do, by design, is run the pricing logic and move the money.
That is where Solvimon comes in. Solvimon is a billing platform built for the most complicated pricing in the market, and the company runs billing today for fintechs, banks, and telcos. However, the AI cloud may be the most complex billing environment yet, and it is precisely what Solvimon was designed for.
The two platforms interface cleanly through APIs, with no custom integration to build. Rafay meters what was consumed and exposes it through the control plane. Solvimon pulls that data, applies the operator's pricing and rate cards, handles invoicing, settlement, and collection, then pushes rate cards back into Rafay so entitlements and spend caps stay in sync. Rafay supplies the metered truth. Solvimon turns it into revenue. For the operator, the path from GPU capacity to a live, billable AI service simply becomes a configuration exercise.
Start metering now
For many workloads, the destination is outcome-based pricing. When an agent completes a piece of work, the customer wants to pay for the work, and private and sovereign operators are best placed for that because their customers are buying finished things: a processed permit, a cleared claim, a reviewed case file.
Every organization standing up its own AI cloud is doing it to control its data, its cost, and its future. Monetization is what makes that control sustainable, turning infrastructure they own into a business they run.
The next generation of AI infrastructure won't differentiate itself only by faster GPUs. It will differentiate itself by how intelligently it packages, measures, governs, and monetizes those resources. The operators who solve that commercial layer earliest may find that pricing becomes as strategic a capability as compute itself.
Rafay and Solvimon are working together to help operators build those capabilities.
Learn more at rafay.co and solvimon.com