Careers

Principal Solutions Architect - U.S.

Full Time
U.S.

About the Role

Rafay seeks a Principal Solutions Architect to enable enterprise customers/Neo clouds  in deploying, operating, and scaling AI/ML workloads on our GPU Platform-as-a-Service offering. This is a hybrid technical leadership and people-management role: alongside hands-on, customer-facing architecture work, this person will build, lead, and grow a team of Solutions Architects, collaborating with platform engineering, MLOps, data science, and infrastructure teams to architect production-ready AI infrastructure solutions built on Kubernetes and GPU-accelerated environments.

Key Responsibilities

Team Leadership & People Management

  • Recruit, hire, and onboard Solutions Architects as the team scales
  • Directly manage a team of Solutions Architects, including workload allocation, coaching, and day-to-day support
  • Set individual and team goals; conduct regular 1:1s and performance reviews
  • Own career development planning for direct reports, including skills growth, promotion readiness, and succession planning
  • Foster an inclusive, high-performing team culture aligned with Rafay's values
  • Manage team capacity, prioritization, and staffing against customer and project demand
  • Partner with sales, engineering, and executive leadership on hiring plans and team structure
  • Mentor and upskill both direct reports and junior team members across the broader organization

Technical & Customer-Facing Responsibilities

  • Design comprehensive AI/ML platform architectures covering inference, training, and data pipelines
  • Develop reference architectures for GPU cluster deployment and LLM serving infrastructure
  • Evaluate inference serving frameworks including vLLM, TGI, and Triton
  • Advise on GPU fabric topology options for distributed training scenarios
  • Design observability strategies using DCGM, OpenTelemetry, and eBPF
  • Translate infrastructure requirements into actionable platform designs
  • Deliver technical presentations, workshops, and proof-of-concept engagements
  • Serve as trusted advisor on AI infrastructure strategy, cost optimization, and scaling
  • Partner with customer stakeholders to understand workload requirements
  • Architect networking, identity management, observability, and security integrations
  • Monitor and troubleshoot production environments for GPU utilization and cluster health
  • Lead root cause analysis for complex customer issues
  • Document reference architectures and implementation best practices

Required Qualifications

  • 8+ years in infrastructure, platform, or solutions engineering roles
  • 3+ years focused on AI/ML infrastructure or MLOps
  • 2+ years of direct people-management experience, including hiring, performance management, and career development of technical staff
  • Demonstrated ability to lead and grow a technical team while remaining hands-on with customers and architecture
  • Deep Kubernetes expertise including cluster lifecycle and RBAC
  • Hands-on experience with NVIDIA GPU infrastructure (H100/H200/B200 preferred)
  • Proficiency with distributed training concepts (NCCL, tensor parallelism)
  • Experience with LLM inference serving and optimization
  • Familiarity with GPU Operator, MIG, SR-IOV, and network fabrics
  • Strong scripting and automation skills (Python, Bash, Go preferred)
  • Ability to communicate complex technical concepts to diverse audiences, including executive stakeholders
  • Experience with AWS, Azure, or GCP platforms
  • Familiarity with monitoring tools like Prometheus, Grafana, and OpenTelemetry
  • Understanding of GPU-based workloads and model serving
  • Proven troubleshooting capabilities for infrastructure issues
  • Excellent communication, coaching, and customer-facing skills

Preferred Qualifications

  • Experience building a Solutions Architecture or technical pre-sales team from the ground up
  • Formal people-management training or leadership certification
  • Enterprise customer support experience in cloud-native environments
  • Familiarity with PyTorch and TensorFlow frameworks
  • Experience with Run:AI and Slurm
  • GPU scheduling and autoscaling expertise
  • Multi-tenant Kubernetes environment knowledge
  • MLOps platform experience
  • Technical workshop leadership experience
  • Relevant certifications (CKA, CKAD, AWS/Azure/GCP Solutions Architect)
  • Understanding of multi-tenant GPU isolation technologies

Why Join Rafay?

Rafay is at the forefront of GPU PaaS technologies and Kubernetes and we offer unique opportunities to join a winning team working on foundational technology for cloud and AI/ML services and enterprises. We work in a collaborative environment that rewards creative thinking and provides opportunities to advance professional careers in advanced technology development. On top of this we offer a fun and dynamic work environment, a competitive salary, robust benefits and attractive stock options. As the first of our kind, we are truly in a class of our own.

Max file size 10MB.
Uploading...
fileuploaded.jpg
Upload failed. Max size for files is 10 MB.
Your application has been successfully submitted.
Oops! Something went wrong while submitting the form.