Careers

Sr. Principal Engineer — Platform -

Full Time

We are looking for a Sr. Principal Engineer to provide technical leadership and make significant contributions to the architecture and development of Rafay's multi-tenant cloud and AI infrastructure platform.

Rafay operates at the intersection of distributed systems, Kubernetes, virtualization, infrastructure automation, observability, and security. This role provides an opportunity to build foundational technologies used to operate complex cloud and AI infrastructure environments at scale.

As a Sr. Principal Engineer, you will serve as a technical leader and hands-on architect, setting technical direction for critical platform capabilities and multiplying the effectiveness of engineering teams across the organization.

You will be expected to move comfortably between architecture and implementation, reason about complex distributed systems and failure modes, and design platforms that are scalable, highly available, observable, secure, and suitable for enterprise and regulated environments.

Responsibilities

  • Design and implement core architectural components for critical services within a large-scale, multi-tenant distributed platform.
  • Set the technical vision and long-term architectural direction for major platform areas and drive alignment across engineering teams.
  • Design highly modular, scalable, resilient, secure, and maintainable distributed services.
  • Architect platform capabilities spanning cloud infrastructure, Kubernetes, virtualization, observability, security, and automation.
  • Drive the design of modern observability capabilities covering metrics, logs, traces, events, infrastructure telemetry, and service health.
  • Develop approaches for correlating information across multiple infrastructure and application layers to improve troubleshooting and operational reliability.
  • Design service health monitoring, alerting, SLI/SLO frameworks, synthetic monitoring, and proactive validation capabilities.
  • Develop capabilities that improve incident detection, troubleshooting, root-cause analysis, and operational automation.
  • Help establish architectures for safe, controlled, and auditable automation of infrastructure operations.
  • Define and drive platform security architecture, including identity, authorization, secrets management, workload isolation, network security, secure APIs, and privileged operations.
  • Design systems suitable for high-security and regulated environments, including government and enterprise deployments.
  • Partner with security and compliance teams to translate compliance requirements into scalable platform capabilities.
  • Contribute to architectures involving confidential computing, trusted execution environments, hardware-backed security, attestation, and protection of sensitive workloads and data.
  • Participate in security architecture reviews and threat modeling for distributed and multi-tenant systems.
  • Ensure systems provide strong auditability for administrative actions, configuration changes, and automated operations.
  • Perform R&D and feasibility analysis on emerging technologies related to cloud infrastructure, AI infrastructure, observability, distributed systems, and security.
  • Assist engineering and operations teams with diagnosing complex production issues and drive systemic improvements based on lessons learned.
  • Lead architecture reviews, design reviews, and code reviews and help establish engineering standards across teams.
  • Mentor engineers and raise the technical bar across the engineering organization.
  • Collaborate with engineering, product management, security, SRE, QA, and customer-facing organizations.
  • Champion a customer-focused engineering culture where production experience informs improvements in reliability, usability, security, and automation.

Skills and Qualifications

  • 12+ years of experience designing, building, and delivering large-scale enterprise software platforms.
  • Demonstrated experience operating as a senior technical architect or Principal-level engineer responsible for significant platform architecture decisions.
  • Deep understanding of distributed systems fundamentals, including:
    • Scalability
    • High availability
    • Resiliency
    • Distributed state
    • Concurrency
    • Failure handling
    • Performance
  • Expert knowledge of one or more programming languages, preferably:
    • Golang
    • Python
  • Strong experience designing and developing microservices and distributed control-plane systems.
  • Excellent troubleshooting and debugging skills across complex production environments.
  • Hands-on experience building services on public cloud platforms such as AWS, Azure, or GCP.
  • Strong practical knowledge of networking fundamentals and protocols including TCP/IP, HTTP/HTTPS, DNS, TLS, load balancing, and network security.
  • Strong experience with Kubernetes and cloud-native architectures.
  • Experience with container orchestration, distributed control planes, APIs, lifecycle management, and multi-tenancy.

Observability & Reliability

  • Deep experience designing or operating large-scale observability and telemetry systems.
  • Experience with metrics, logging, distributed tracing, event processing, and infrastructure monitoring.
  • Hands-on experience with technologies such as OpenTelemetry, Prometheus, Grafana, or comparable platforms.
  • Experience defining and operating:
    • SLIs
    • SLOs
    • Service-health models
    • Alerting strategies
    • Capacity and performance monitoring
  • Experience designing systems for proactive problem detection, telemetry correlation, and production troubleshooting.
  • Experience with synthetic monitoring, anomaly detection, or automated service validation is highly desirable.
  • Experience applying AI or automation to improve operational diagnosis and reliability is a plus.

Security

  • Strong understanding of cloud-native and distributed-systems security.
  • Hands-on experience with areas such as:
    • IAM and RBAC
    • Least privilege
    • Workload identity
    • PKI and certificate lifecycle
    • Secrets management
    • TLS/mTLS
    • Network security and segmentation
    • API security
    • Secure software supply chains
    • Vulnerability management
    • Security logging and auditing
  • Experience designing secure multi-tenant platforms with strong workload and tenant isolation.
  • Experience performing threat modeling and security architecture reviews.
  • Familiarity with policy-as-code, automated security controls, and security telemetry is desirable.

FedRAMP & Regulated Environments

  • Experience designing, implementing, or operating software platforms in FedRAMP, U.S. government, or similarly regulated environments is highly desirable.
  • Experience translating regulatory and security requirements into technical architecture and software controls.
  • Experience with continuous compliance, security monitoring, audit evidence, vulnerability management, and configuration governance is desirable.

Confidential Computing

  • Understanding of confidential computing and trusted execution environments.
  • Familiarity with hardware-backed security technologies available across modern CPU, GPU, virtualization, and cloud platforms.
  • Understanding of concepts such as:
    • Hardware roots of trust
    • Secure and measured boot
    • Attestation
    • Trusted execution environments
    • Memory encryption
    • Data-in-use protection
  • Experience integrating confidential-computing technologies into Kubernetes, virtualization, or cloud platforms is highly desirable.

Technical Leadership

  • Proven ability to make architecture decisions involving substantial scale, reliability, security, and operational complexity.
  • Demonstrated ability to influence technical direction and build consensus across multiple teams.
  • Ability to identify architectural limitations and drive longer-term improvements across multiple releases.
  • Experience solving difficult customer production problems and converting those learnings into platform improvements.
  • Strong ability to mentor senior engineers and establish engineering practices that scale beyond an individual team.

Preferred Background

A particularly strong candidate will have experience across several of the following areas:

Distributed Systems | Kubernetes | Cloud Infrastructure | Observability | Security | GPU/AI Infrastructure | Virtualization | FedRAMP | Confidential Computing

Candidates are not expected to be experts in every area. We are looking for individuals with deep expertise in several domains and sufficient architectural breadth to reason across the entire platform.

WHY JOIN RAFAY

Rafay is building foundational infrastructure for modern cloud and AI environments. Engineers at Rafay work on challenging problems involving distributed systems, Kubernetes, virtualization, observability, security, and infrastructure automation.

We offer an environment where senior engineers can influence platform architecture, work on technically demanding infrastructure problems, and help define the next generation of cloud and AI infrastructure management.

Max file size 10MB.
Uploading...
fileuploaded.jpg
Upload failed. Max size for files is 10 MB.
Your application has been successfully submitted.
Oops! Something went wrong while submitting the form.