Operationalizing AI Fabrics with Aviz ONES, NVIDIA Spectrum-X, and Rafay
Discover the new AI operations model available to enterprises that enables self-service consumption and cloud-native orchestration for developers.
Rafay-powered SLURM as a Service delivers fully managed, multi-tenant SLURM environments for high-performance computing workloads as a cloud-like, on-demand service. Through automated, BCM-based cluster bring-up with secure per-tenant separation and governance built in, tenants request a cluster and Rafay provisions, schedules, and governs it.
.webp)
Slurm as a Service delivers the Slurm workload manager through a fully managed, on-demand platform. Instead of manually deploying and maintaining HPC clusters, organizations can provision per-tenant Slurm clusters, submit jobs to managed queues, and allocate compute resources through self-service, while the platform automates both the underlying Kubernetes cluster's provisioning and the Slurm environment's scheduling, governance, and lifecycle management.
Rafay automates bring-up of the underlying Kubernetes cluster and layers Slurm on top via the open-source Slinky Slurm Operator, enabling providers and enterprises to deliver HPC resources as scalable, self-service Slurm clusters.
Familiar HPC scheduling, delivered as a governed, on-demand service.
Automated BCM-based bring-up removes manual management of the underlying Kubernetes cluster.
Login nodes let users submit jobs immediately
Multiple tenants or teams with full isolation and operational control
Provisioning, scaling, and teardown handled automatically
Provision fully managed Slurm clusters on demand through a self-service portal or API.
Automate cluster provisioning, scaling, lifecycle management, and governance to eliminate manual administration.
Run HPC and AI workloads on shared CPU and GPU infrastructure to improve resource efficiency.
Deliver isolated, multi-tenant Slurm environments with centralized governance and policy controls.
Turn HPC infrastructure into a managed, consumption-based service with built-in usage metering and chargeback.
Deliver production-ready HPC environments that support research, engineering, and AI workloads from a single platform.
Whether you're delivering HPC services to customers or supporting internal research and engineering teams, Rafay helps simplify HPC operations while maximizing infrastructure utilization.
Find answers to common questions about our SLURM as a service offering below.
Slurm (Simple Linux Utility for Resource Management) is an open-source workload manager used to schedule and manage compute-intensive workloads across high-performance computing (HPC) clusters. It allows users to submit jobs, request compute resources, join queues, and run workloads across CPUs, GPUs, and other infrastructure.
Originally designed for HPC environments, Slurm is widely used in research, engineering, scientific computing, and AI/ML workloads where teams need efficient resource allocation, scheduling, and cluster management.
Slurm as a Service enables cloud providers, neoclouds, research organizations, and enterprises to offer managed HPC and GPU compute environments without building and operating the entire platform themselves.
Common users include:
Managing Slurm independently requires significant expertise to deploy, configure, secure, scale, and maintain HPC clusters. Rafay automates bring-up of the underlying Kubernetes cluster and simplifies delivery of per-tenant Slurm clusters on top of it, enabling self-service access and enterprise-grade governance.
Yes. Rafay supports provisioning and lifecycle management across bare metal, virtual machines, Kubernetes, and SLURM environments, allowing providers to deliver standardized infrastructure and AI services from a single platform.
Yes. Through the open-source Slinky Slurm Operator, Slurm's scheduler runs on top of the same Kubernetes cluster that Rafay provisions, so Slurm-based HPC jobs and native Kubernetes workloads share the same underlying infrastructure rather than running as separate stacks.
Slurm is commonly used for batch-oriented HPC workloads, including large-scale simulations, scientific computing, and AI training jobs. Kubernetes is often used for containerized applications, AI inference, and cloud-native services.
Using both enables organizations to support a broader range of workloads while maximizing GPU utilization across their infrastructure.
Yes, Rafay provides infrastructure orchestration and workflow automation for cloud-native (Kubernetes) and AI use cases for enterprises, cloud providers, neoclouds, and Sovereign AI clouds. Rafay helps companies deploy a Platform-as-a-Service (PaaS) experience that supports both CPU-only and GPU-accelerated compute environments. Platform teams can quickly set up and deliver customized self-service experiences for developers and data scientists, typically within days or weeks. This flexible platform allows end-users to easily access the computational resources they need, whether it’s standard CPU processing or more powerful GPU capabilities. Rafay’s solution streamlines the deployment and management of diverse computing environments, making it easier for organizations to support a wide range of applications, from standard software to complex AI/ML projects.
Talk with Rafay experts to assess your infrastructure, explore your use cases, and see how teams like yours operationalize AI/ML and cloud-native initiatives with self-service and governance built in.