Inference Optimization Engineer (Sr Engineer / Sta5 Engineer / Principal Engineer) - India
Role Overview
This engineering role begins where model training ends. Once a machine learning or large language model (LLM) is trained, an inference optimization engineer figures out how to serve it efficiently to real-
world users under strict speed and cost limits. The position sits between research and low-level systems engineering. It requires balancing speed—measured as Time-to-First-Token (TTFT) and tokens per second—with output accuracy and hardware costs.
Key Responsibilities
• Model Compression: Apply techniques like quantization (FP8, INT8, INT4), pruning, and distillation to shrink model size without losing accuracy.
• Runtime & Serving Tuning: Configure and optimize high-throughput serving frameworks such as vLLM, SGLang, TensorRT-LLM, or Triton Inference Server.
• Memory & Caching Management: Optimize KV cache usage, continuous batching, and speculative decoding to handle long context windows and multiple concurrent users.
• Profiling & Benchmarking: Use tools like Nsight Systems, PyTorch Profiler, and Triton Metrics to identify memory bandwidth or compute bottlenecks on GPU/TPU clusters.
• Hardware Co-Optimization: Tune custom kernels (CUDA/Triton) and manage distributed setups across multi-GPU clusters using tensor and pipeline parallelism
Key Requirements & Qualifications
• Experience: 4+ years in systems programming, high-performance computing (HPC), or ML infrastructure.
• Core Languages: Strong proficiency in Python or Go, with familiarity reading or writing CUDA/Triton kernels.
• Frameworks: Hands-on production experience with modern inference engines (vLLM, TensorRT-LLM, SGLang) and profiling tools (Nsight).
• Foundational Knowledge: Deep understanding of transformer architectures, memory hierarchies, arithmetic intensity, and roofline performance models.










