Certifying Tenant GPU Kubernetes Clusters with NVIDIA Cluster Readiness Engine
Learn how NeoClouds can use NVIDIA Cluster Readiness Engine to certify GPU Kubernetes clusters before tenant handoff and validate AI workload readiness.
Read Now

This blog was co-authored by Rupinder (Robbie) Gill and Haseeb Budhani.
In a previous blog, we discussed reasons why companies are choosing to operate many small clusters.As we’ve been following up with engineers we met at Kubecon 2019, we are finding that a high percentage of teams are running two or more Kubernetes distributions across their public cloud and on-premise footprints. The common belief seems to be that cloud providers will do their best to optimize their Kubernetes offerings for their infrastructure, so its best to use EKS in AWS, GKE in GCP, AKS in Azure, and OpenShift or PKS on premises.At Rafay, we follow the same methodology: We use the resident managed Kubernetes service in our cloud provider of choice instead of spinning up, for example, v1.16.1 ourselves on virtual machines. And - of course - we run many, small clusters.There are many hurdles to cross in keeping modern applications operational, and if the public cloud (or VMware & RedHat) is able to take away the pain of keeping the Kubernetes control plane up and running, why would anyone not leverage that? What’s more, the cost of running Kubernetes in public clouds is fast approaching zero. You pay for the worker node VMs and the master node costs are a rounding error or downright free.Operating services across multiple clusters and multiple distros simplifies the development process. Ongoing complexity is reduced because developers no longer have to add complex logic in each service to address environmental characteristics such as service meshes, storage classes and admission controllers. With a growing cluster fleet that may span multiple clouds & data centers, and leveraging multiple Kubernetes distributions, SRE/Ops need tooling to manage their cluster fleet. SRE/Ops teams must now solve for:
We live in a multi-distro, multi-cluster world. But this is an issue that cloud providers don’t have an incentive to solve for you. When engaging with a vendor focused on helping your SRE/Ops teams address the complexity of managing multiple clusters running a variety of distros, be sure to ask for their plan for the above 3 requirements. Rafay delivers these capabilities today and can simplify ongoing operations for Kubernetes environments as a service. Feel free to get in touch to learn more.

Learn how NeoClouds can use NVIDIA Cluster Readiness Engine to certify GPU Kubernetes clusters before tenant handoff and validate AI workload readiness.
Read Now
Rafay CEO Haseeb Budhani joins theCUBE to discuss sovereign AI, multi-tenant cloud platforms, GPU monetization, and the shift from infrastructure to AI services.
Read Now

Over the past several years we’ve experienced a tremendous amount of change in the Kubernetes management and container orchestration market. Years ago, Kubernetes was used to support a relatively small number of clusters in lab environments, handling mostly corner use cases, and seen as a simple cluster management tool that was used by DevOps and IT Ops.
Read Now