Rafay and VotalAI Partner to Deploy AI Securely at Scale
.png)
New integration brings runtime AI guardrails and agentic security to Rafay-powered AI services, helping NeoClouds and enterprises protect models, agents, tools, and customer interactions.
Rafay and VotalAI are partnering to help enterprises, NeoClouds, and sovereign AI providers deploy AI services with security and policy enforcement embedded directly into the runtime.
The collaboration combines Rafay’s platform for operating, governing, and monetizing AI infrastructure with VotalAI’s runtime security platform for enterprise AI applications and agents. Together, the companies are addressing a growing operational challenge: as organizations expose models through APIs and connect agents to enterprise tools and data, traditional infrastructure security no longer covers the full AI interaction path.
Rafay makes accelerated infrastructure consumable through governed, multi-tenant services, including standardized inference APIs and token-metered consumption through Rafay Token Factory. VotalAI adds an inline security and policy layer that inspects AI interactions and enforces controls across prompts, model responses, agent identities, and tool calls.
The result is a complementary architecture: Rafay deploys and operates enterprise AI services. VotalAI helps secure how those services behave at runtime.

Closing the Runtime Security Gap
AI infrastructure security has traditionally concentrated on the systems surrounding the model: identity, network isolation, access controls, workload governance, and data residency. Those controls remain essential, but they cannot independently determine whether a prompt is attempting to manipulate a model, whether an agent should be permitted to call a particular tool, or whether a generated response violates a customer-specific policy.
These risks become more pronounced as organizations advance from standalone model inference to agentic applications. An AI agent may reason across multiple steps, retrieve enterprise data, invoke an MCP server, call third-party services, and take actions on behalf of a user. Every transition introduces another policy decision and another potential attack surface.
Runtime authorization depends on a question infrastructure identity does not answer: which agent is making a request, and on whose authority. API keys identify a tenant for metering, and enterprise identity providers identify the human who authenticated to an application. Neither establishes the provenance of the agent itself or the lineage of a request as it passes between agents, tools, and services. VotalAI attests each agent at issuance, binds short-lived credentials to that provenance, and preserves delegation lineage across every hop. The Rafay gateway verifies those claims offline, without adding an identity round trip.
VotalAI operates inline with these interactions. Its platform is designed to block prompt injection, detect sensitive-data exposure, restrict topics, enforce role-based tool authorization, and record runtime decisions for audit and investigation. Votal supports model calls and agentic workflows across frameworks and environments rather than requiring customers to adopt a single model provider or application stack.
The partnership addresses two important security requirements:
- Operator-level attack and manipulation protection: Inspecting traffic and enforcing controls before malicious or unauthorized requests reach models, tools, or downstream systems.
- Customer-specific runtime guardrails: Allowing operators to apply distinct policies for each enterprise, tenant, application, or regulated workload.
Security Across Models, MCP Servers, and AI Agents
The integration is being developed around LiteLLM, enabling VotalAI to serve as a drop-in security layer within OpenAI-compatible inference workflows. This approach gives operators a pragmatic starting point without requiring them to redesign applications or replace their existing model-serving architecture.
The joint solution spans three principal runtime layers:
- Model inference: Inspect prompts and responses, detect adversarial activity, redact sensitive information, enforce content policies, and maintain tenant-specific audit trails.
- MCP server environments: Validate MCP servers, govern tool access, restrict sensitive operations, and apply data-handling policies when agents interact with external systems.
- Agentic applications: Protect multi-step workflows built with frameworks such as LangChain, CrewAI, OpenAI Agents SDK, and custom orchestration systems.
VotalAI provides enforcement for model calls, tool invocations, and multi-tenant environments, with support for declarative configuration, policy versioning, audit records, and runtime actions including blocking, redaction, and transformation.
These policies are not static. VotalAI continuously tests live model and agent endpoints against a library of more than 140 adversarial manipulation strategies, and findings from that testing are used to retrain the guardrail model, so enforcement tracks attacker technique rather than lagging it.
Because guardrails sit in the request path, performance is a critical design constraint. The companies are prioritizing low-latency enforcement so security does not materially degrade interactive inference or agent execution. Broad multilingual coverage is also an important requirement for NeoClouds and sovereign providers serving diverse regional markets.
Making Hyperscaler-Grade Guardrails Accessible
Major public cloud and model platforms increasingly surround their AI services with safety filters, policy controls, and governance tooling. Regional providers and enterprises building on their own infrastructure need comparable protections, but often lack the security engineering resources required to assemble and continuously maintain them.
The Rafay and VotalAI partnership is intended to make those capabilities accessible within privately operated, regional, and sovereign AI platforms.
This is particularly relevant for NeoClouds responding to sovereignty-driven demand. Their customers may require models and data to remain within a specific country or jurisdiction, but infrastructure locality alone does not resolve application-layer security. Providers must also control how models are accessed, what information they can return, and which systems an agent can operate.
Sovereignty carries a second-order requirement that is easily overlooked. A guardrail that classifies prompts by calling a frontier model API sends every inspected prompt outside the operator’s jurisdiction, reintroducing the residency problem the deployment was intended to solve. There is a reliability dimension as well: frontier models apply their own safety policies to the content they are asked to evaluate, and may decline to classify an adversarial payload at all. VotalAI’s guardrail model is open-weight and fine-tuned on a proprietary attack corpus, and operates on the operator’s own accelerators alongside NVIDIA Nemotron Safety 3.5. Classification remains within the same boundary as inference.
Industries such as financial services, healthcare, telecommunications, government, defense, and critical infrastructure can apply differentiated policies without forcing every customer into the same risk model. A healthcare tenant may require aggressive protection of personal and clinical information, while a financial-services tenant may emphasize tool authorization, transaction controls, and comprehensive auditability.

From Token Factory to a Secure AI Services Platform
Rafay Token Factory enables operators to expose open-source and enterprise AI models through standardized APIs, meter consumption at the token level, and deliver services across shared GPU infrastructure with multi-tenancy and governance built in. This allows providers to progress beyond selling raw GPU capacity toward delivering AI services aligned with customer usage.
VotalAI extends that operating model with runtime enforcement across the interaction lifecycle.
Together, the platforms are designed to help customers:
- Launch protected AI services faster without building a proprietary guardrail and agent-security stack.
- Apply policies by tenant or use case across prompts, outputs, tools, agents, and sensitive data.
- Support sovereign AI requirements through deployable controls that can operate alongside regional, private, or air-gapped infrastructure.
- Expand AI services confidently from basic inference APIs into MCP-enabled and agentic applications.
A Shared Foundation for Secure Enterprise AI
Mohan Atreya, Chief Product Officer, Rafay Systems
“Organizations are moving rapidly from GPU infrastructure to consumable AI services. Our work with VotalAI helps ensure that security and runtime policy enforcement advance alongside that transition, giving operators a stronger foundation for serving enterprise and sovereign customers.”
Jyotirmoy Sundi, Chief Technology Officer, VotalAI“Enterprise AI security must extend beyond the model endpoint. By combining VotalAI’s runtime policy plane with Rafay’s AI infrastructure and Token Factory capabilities, we can help customers govern every interaction, from the initial prompt through agent actions and model output.”
The initial LiteLLM-based integration will be made available shortly. Additional platform-level integrations are planned to deepen policy management, telemetry, tenant administration, and enforcement across Rafay-powered AI environments.
As AI systems become more autonomous, protecting infrastructure will remain necessary but insufficient. Models, agents, tools, and data must be governed as a connected runtime system.
Rafay and VotalAI are building that system for organizations that need to deploy enterprise AI securely, operate it across multiple tenants, and scale it without surrendering performance, control, or sovereignty.
Learn more about the Rafay Token Factory and VotalAI.








