About the role
We serve inference at cost-per-token margins that leave no room for a sloppy stack. A few percent of throughput is the difference between a healthy unit economic and an unprofitable one, so the serving layer gets treated as a first-class engineering surface rather than a deployment detail.
You will own that layer. That means choosing and tuning the right serving runtime per workload, owning the quantisation and evaluation work that proves a trade-off is safe, and maintaining the benchmarking discipline that keeps our published numbers honest. You will work across the GPU systems team below you and the platform and product teams above, and your results show up directly in what customers pay and what latency they see.
What success looks like
- Latency and throughput for our top workloads improve measurably, with benchmarks to evidence it.
- Every quantisation or runtime change ships with an evaluation that shows what it cost in quality.
- Benchmark runs are reproducible by someone who did not write them.
Responsibilities
- Own the model serving stack end to end, including runtime selection, configuration and upgrade path for each workload class.
- Tune throughput and latency through continuous batching, KV cache strategy, paged attention, tensor and pipeline parallelism, and speculative decoding where it pays.
- Run quantisation work such as FP8, AWQ and GPTQ, and build the evaluation harness that quantifies the quality trade-off rather than guessing at it.
- Maintain reproducible benchmarking across model families and hardware generations, including the methodology write-ups behind published results.
- Profile and eliminate bottlenecks across the whole request path, covering GPU kernels, host-side overhead, batching and network.
- Partner with the GPU systems team on node-level tuning, and with platform engineering on autoscaling and scheduling behaviour.
- Keep a clear eye on serving-side cost per token, and make the trade-offs explicit when latency and cost pull in opposite directions.
Requirements
- Shipped and operated at least one production inference stack on modern datacentre GPUs such as H100 or A100.
- Working depth in one or more serving runtimes, for example vLLM, TensorRT-LLM, SGLang or Triton Inference Server.
- Solid understanding of transformer inference mechanics: attention, KV cache behaviour, batching strategies and where the memory actually goes.
- Strong Python, and enough comfort in CUDA or C++ to read a kernel and reason about it.
- Real experience with benchmarking and evaluation methodology, including knowing how to avoid fooling yourself with a flattering measurement.
- Familiarity with quantisation and its quality implications, rather than treating it as a switch to flip.
- Ability to communicate performance results to non-specialists without either hand-waving or drowning them.
Nice to have
- Open-source contributions to a serving or inference project.
- Experience with fine-tuning workloads as well as inference.
- Exposure to multi-modal or embedding workloads alongside text generation.
- Familiarity with public inference benchmarking suites and their failure modes.
Role reference: YJ-00003