Careers

Senior Solutions Engineer - U.S.

Full Time
U.S.

About the Role

Rafay seeks a customer-focused, technically skilled Senior Solutions Engineer.. The role involves partnering with enterprise customers/Neo clouds to architect, deploy, and optimize cloud-native infrastructure using the Rafay platform and Kubernetes ecosystems, including GPU-accelerated environments for AI/ML and inference workloads.

Key Responsibilities

Core Implementation Responsibilities

  • Serve as primary technical lead for customer implementation engagements
  • Partner with customers on requirements gathering, architecture reviews, and solution design
  • Design and configure Rafay platform capabilities across public, private, and hybrid cloud environments
  • Troubleshoot complex infrastructure, networking, Kubernetes, and virtualization issues
  • Collaborate with Customer Success, Product, and Engineering teams on technical challenges
  • Manage customer issues within SLAs while maintaining expectations
  • Reproduce and analyze customer-reported issues; communicate findings to internal teams
  • Develop technical documentation, implementation guides, and runbooks
  • Mentor junior engineers and contribute to process improvements
  • Stay current on Rafay platform capabilities and emerging industry trends

GPU & AI/ML Infrastructure

  • Deploy and configure GPU-accelerated Kubernetes clusters for AI/ML workloads, including GPU Operator, MIG, and SR-IOV setup
  • Implement and validate LLM inference serving stacks (vLLM, TGI, Triton) during customer implementations
  • Troubleshoot GPU cluster health, utilization, and GPU fabric/networking issues (NCCL, RoCEv2)
  • Configure GPU observability using DCGM, OpenTelemetry, and related monitoring tooling
  • Advise customers on distributed training and inference topology options for NVIDIA GPU infrastructure (H100/H200/B200)

Minimum Qualifications

  • 10+ years of experience in customer-facing technical roles, including implementation or consulting
  • Hands-on AI/ML engineering experience with model inference and deployment workflows
  • Hands-on experience with NVIDIA GPU infrastructure (H100/H200/B200 preferred)
  • Familiarity with GPU Operator, MIG, SR-IOV, and GPU fabric/networking fundamentals
  • Experience with LLM inference serving frameworks such as vLLM, TGI, or Triton
  • Strong Kubernetes and cloud-native environment expertise
  • Excellent written and verbal communication abilities
  • Deep troubleshooting skills across networking, virtualization, container orchestration, and public cloud platforms
  • Infrastructure automation and IaC (Terraform preferred) experience desired
  • Linux systems administration and distributed systems knowledge
  • Proven ability to lead technical projects and manage multiple engagements
  • Bachelor's degree in Computer Science, Engineering, IT, or equivalent experience

Preferred Qualifications

  • Experience with Run:AI and Slurm for GPU scheduling
  • GPU autoscaling and multi-tenant GPU isolation expertise
  • Familiarity with DCGM-based GPU observability and Prometheus/Grafana dashboards
  • Relevant certifications (CKA, CKAD, or NVIDIA certifications)

Why Join Rafay?

Rafay is at the forefront of GPU PaaS technologies and Kubernetes and we offer unique opportunities to join a winning team working on foundational technology for cloud and AI/ML services and enterprises. We work in a collaborative environment that rewards creative thinking and provides opportunities to advance professional careers in advanced technology development. On top of this we offer a fun and dynamic work environment, a competitive salary, robust benefits and attractive stock options. As the first of our kind, we are truly in a class of our own. 

Max file size 10MB.
Uploading...
fileuploaded.jpg
Upload failed. Max size for files is 10 MB.
Your application has been successfully submitted.
Oops! Something went wrong while submitting the form.