Senior Solutions Engineer - U.S.
About the Role
Rafay seeks a customer-focused, technically skilled Senior Solutions Engineer.. The role involves partnering with enterprise customers/Neo clouds to architect, deploy, and optimize cloud-native infrastructure using the Rafay platform and Kubernetes ecosystems, including GPU-accelerated environments for AI/ML and inference workloads.
Key Responsibilities
Core Implementation Responsibilities
- Serve as primary technical lead for customer implementation engagements
- Partner with customers on requirements gathering, architecture reviews, and solution design
- Design and configure Rafay platform capabilities across public, private, and hybrid cloud environments
- Troubleshoot complex infrastructure, networking, Kubernetes, and virtualization issues
- Collaborate with Customer Success, Product, and Engineering teams on technical challenges
- Manage customer issues within SLAs while maintaining expectations
- Reproduce and analyze customer-reported issues; communicate findings to internal teams
- Develop technical documentation, implementation guides, and runbooks
- Mentor junior engineers and contribute to process improvements
- Stay current on Rafay platform capabilities and emerging industry trends
GPU & AI/ML Infrastructure
- Deploy and configure GPU-accelerated Kubernetes clusters for AI/ML workloads, including GPU Operator, MIG, and SR-IOV setup
- Implement and validate LLM inference serving stacks (vLLM, TGI, Triton) during customer implementations
- Troubleshoot GPU cluster health, utilization, and GPU fabric/networking issues (NCCL, RoCEv2)
- Configure GPU observability using DCGM, OpenTelemetry, and related monitoring tooling
- Advise customers on distributed training and inference topology options for NVIDIA GPU infrastructure (H100/H200/B200)
Minimum Qualifications
- 10+ years of experience in customer-facing technical roles, including implementation or consulting
- Hands-on AI/ML engineering experience with model inference and deployment workflows
- Hands-on experience with NVIDIA GPU infrastructure (H100/H200/B200 preferred)
- Familiarity with GPU Operator, MIG, SR-IOV, and GPU fabric/networking fundamentals
- Experience with LLM inference serving frameworks such as vLLM, TGI, or Triton
- Strong Kubernetes and cloud-native environment expertise
- Excellent written and verbal communication abilities
- Deep troubleshooting skills across networking, virtualization, container orchestration, and public cloud platforms
- Infrastructure automation and IaC (Terraform preferred) experience desired
- Linux systems administration and distributed systems knowledge
- Proven ability to lead technical projects and manage multiple engagements
- Bachelor's degree in Computer Science, Engineering, IT, or equivalent experience
Preferred Qualifications
- Experience with Run:AI and Slurm for GPU scheduling
- GPU autoscaling and multi-tenant GPU isolation expertise
- Familiarity with DCGM-based GPU observability and Prometheus/Grafana dashboards
- Relevant certifications (CKA, CKAD, or NVIDIA certifications)
Why Join Rafay?
Rafay is at the forefront of GPU PaaS technologies and Kubernetes and we offer unique opportunities to join a winning team working on foundational technology for cloud and AI/ML services and enterprises. We work in a collaborative environment that rewards creative thinking and provides opportunities to advance professional careers in advanced technology development. On top of this we offer a fun and dynamic work environment, a competitive salary, robust benefits and attractive stock options. As the first of our kind, we are truly in a class of our own.










