Careers

Technical Solutions Architect – GPU Platform - United States (SF Bay Area Preferred, Remote OK)

Full Time
United States (SF Bay Area Preferred, Remote OK)

The Role

We are seeking a Technical Solutions Architect – Rafay GPU Platform to join our product team. You will work closely with product managers, engineering, technical publications, and field solution architects to define, validate, and document how customers deploy and operate Rafay’s GPU Platform.

This is a very hands-on and technical role focused on turning product capabilities into real-world architectures, deployment patterns, use cases, and troubleshooting guidance. You will build reference environments, validate end-to-end deployment scenarios, document architecture and operational best practices, and help customers understand how to successfully deploy GPU infrastructure and AI workloads using Rafay.

The ideal candidate combines strong infrastructure and cloud-native architecture skills with the ability to communicate complex technical concepts clearly. You should be comfortable working across GPU infrastructure, networking, storage, virtualization, Kubernetes and AI/ML workloads.

You will have the opportunity to work on cutting-edge infrastructure spanning GPUs, AI/ML, Generative AI, Kubernetes, virtualization, and modern data center technologies.

Key Responsibilities

Solution Architecture & Reference Architectures

  • Design and document GPU PaaS reference architectures for cloud providers, enterprises, service providers, and AI infrastructure operators.
  • Develop architecture patterns covering GPU compute, Kubernetes, virtual machines, bare metal, networking, storage, security, observability, and AI workloads.
  • Create detailed architecture diagrams illustrating platform components, infrastructure dependencies, traffic flows, integrations, and deployment models.
  • Build and maintain reproducible reference environments that demonstrate recommended deployment patterns.
  • Define architecture guidance for production considerations including scalability, high availability, multi-tenancy, security, and operational resilience.

Deployment Guides & Implementation Patterns

  • Develop comprehensive, hands-on deployment guides that help customers move from infrastructure prerequisites to production-ready environments.
  • Document infrastructure prerequisites, installation workflows, configuration options, integrations, validation steps, and operational best practices.
  • Create deployment patterns for common GPU PaaS environments, including Kubernetes clusters, GPU VMs, bare-metal GPU servers, AI development environments, and inference services.
  • Develop configuration examples using technologies such as Kubernetes YAML, Helm, Terraform, APIs, CLI tools, and automation frameworks.
  • Validate documented procedures against real environments to ensure they are accurate and reproducible.

GPU PaaS Use Cases

  • Identify and document common customer use cases and translate them into repeatable solution architectures and implementation patterns.
  • Explain architectural tradeoffs and provide guidance on selecting the appropriate deployment model for different workloads.

Troubleshooting & Operational Guidance

  • Develop troubleshooting guides, runbooks, and diagnostic workflows for common deployment and operational issues.
  • Reproduce customer and field issues in reference environments to understand root causes and document resolution procedures.
  • Create troubleshooting decision trees covering infrastructure, Kubernetes, GPU drivers/operators, networking, storage, scheduling, and AI workloads.
  • Work with engineering and support teams to identify recurring issues and convert lessons learned into reusable operational guidance.
  • Document health checks, validation procedures, logs, metrics, and diagnostic commands customers can use to operate environments effectively.

Customer & Field Intelligence

  • Gather feedback from customer deployments, POCs, support cases, and field engagements.
  • Identify recurring architecture patterns, deployment challenges, and operational issues.
  • Translate field learnings into improved reference architectures, deployment guides, troubleshooting content, and product requirements.
  • Partner with product management to identify gaps that can be addressed through product enhancements or improved operational workflows.

Technical Demonstrations & Enablement

  • Build and maintain demo environments showcasing GPU PaaS architectures and use cases.
  • Create technical demonstrations and walkthroughs that explain deployment workflows and architecture concepts.
  • Develop presentations and enablement materials for solution architects, partners, customers, and internal teams.
  • Record technical demos and deployment walkthroughs for training and product launches.
  • Contribute hands-on technical blogs and architecture content highlighting real-world implementation patterns.

Technical Qualifications

  • 3–6+ years of experience in solutions architecture, solutions engineering, infrastructure engineering, DevOps, platform engineering, or a related technical role.
  • Strong understanding of Kubernetes and cloud-native infrastructure.
  • Experience designing or operating infrastructure across public cloud, private cloud, or data center environments.
  • Understanding of GPU infrastructure and AI/ML workloads, including GPU scheduling, drivers, operators, resource allocation, and workload lifecycle.
  • Familiarity with virtualization, bare-metal infrastructure, networking, storage, and infrastructure automation.
  • Experience with technologies such as:
    • Kubernetes, Helm, and container runtimes
    • Terraform and infrastructure-as-code
    • Python, Bash, YAML, and REST APIs
    • GPU infrastructure and NVIDIA software stacks
    • Networking and load balancing
    • Persistent storage and CSI-based storage platforms
    • Observability, metrics, logging, and troubleshooting tools
  • Ability to build environments, deploy workloads, troubleshoot failures, and document the complete process.

Experience with technologies such as Slurm, KubeVirt, NVIDIA GPU Operator, NVIDIA DRA, MIG, vLLM, Ray, Kubeflow, or AI inference platforms is a plus.

Communication & Technical Writing

  • Exceptional written and verbal communication skills for highly technical audiences.
  • Ability to translate complex infrastructure concepts into clear architectures, deployment procedures, and operational guidance.
  • Strong ability to create architecture diagrams and visually communicate system designs.
  • Experience creating technical presentations, demonstrations, deployment guides, or reference architectures.
  • Comfortable explaining both how a solution works and why a particular architecture should be used.

Collaboration & Execution

  • Comfortable working cross-functionally with product management, engineering, support, technical publications, solution architects, partners, and customers.
  • Strong hands-on problem-solving and troubleshooting skills.
  • Ability to independently build and validate technical environments.
  • Proven ability to manage multiple technical deliverables in a fast-moving environment.
  • Self-starter who enjoys exploring new infrastructure technologies and turning that knowledge into practical guidance for customers.

Why Join Us

  • Work at the intersection of GPU infrastructure, Kubernetes, cloud platforms, and AI.
  • Define how customers architect and deploy production GPU PaaS environments.
  • Build reference architectures and deployment patterns that directly influence customer success.
  • Work hands-on with emerging GPU, AI, networking, storage, and cloud-native technologies.
  • Partner closely with engineering and product teams and help shape the evolution of the Rafay platform.

Max file size 10MB.
Uploading...
fileuploaded.jpg
Upload failed. Max size for files is 10 MB.
Your application has been successfully submitted.
Oops! Something went wrong while submitting the form.