Careers

Inference Optimization Engineer (Sr Engineer / Sta5 Engineer / Principal Engineer) - India

Full Time
India

Role Overview

This engineering role begins where model training ends. Once a machine learning or large language model (LLM) is trained, an inference optimization engineer figures out how to serve it efficiently to real-

world users under strict speed and cost limits. The position sits between research and low-level systems engineering. It requires balancing speed—measured as Time-to-First-Token (TTFT) and tokens per second—with output accuracy and hardware costs.

Key Responsibilities

• Model Compression: Apply techniques like quantization (FP8, INT8, INT4), pruning, and distillation to shrink model size without losing accuracy.

• Runtime & Serving Tuning: Configure and optimize high-throughput serving frameworks such as vLLM, SGLang, TensorRT-LLM, or Triton Inference Server.

• Memory & Caching Management: Optimize KV cache usage, continuous batching, and speculative decoding to handle long context windows and multiple concurrent users.

• Profiling & Benchmarking: Use tools like Nsight Systems, PyTorch Profiler, and Triton Metrics to identify memory bandwidth or compute bottlenecks on GPU/TPU clusters.

• Hardware Co-Optimization: Tune custom kernels (CUDA/Triton) and manage distributed setups across multi-GPU clusters using tensor and pipeline parallelism

Key Requirements & Qualifications

• Experience: 4+ years in systems programming, high-performance computing (HPC), or ML infrastructure.

• Core Languages: Strong proficiency in Python or Go, with familiarity reading or writing CUDA/Triton kernels.

• Frameworks: Hands-on production experience with modern inference engines (vLLM, TensorRT-LLM, SGLang) and profiling tools (Nsight).

• Foundational Knowledge: Deep understanding of transformer architectures, memory hierarchies, arithmetic intensity, and roofline performance models.

Max file size 10MB.
Uploading...
fileuploaded.jpg
Upload failed. Max size for files is 10 MB.
Your application has been successfully submitted.
Oops! Something went wrong while submitting the form.