# Senior ML Systems Engineer, Inference at Runpod

- Company: Runpod
- What the company does: AI infrastructure with on-demand GPUs and serverless compute. Run training, inference, and batch workloads on the cloud with Runpod. Backed by AI Grant.
- Company website: https://www.runpod.io/
- Type: Startups (AI role)
- Level: Senior
- Location: Remote - USA
- Work setup: Remote
- Pay: $150K to $220K base salary per year (USD)
- Posted: 2026-09-25
- Apply by: 2026-11-09
- Apply: https://jobs.ashbyhq.com/runpod/6c67ef7b-ae53-42fd-a1ca-6c635f8f00ff
- Page: https://www.1752.vc/careers/jobs/runpod-senior-ml-systems-engineer-inference/

## About the role

Define how we measure inference performance, including throughput, time to first token, inter-token latency, and cost per token, and build the tooling that makes those measurements rigorous and repeatable. Profile and diagnose performance problems across the serving stack, from scheduling and memory management down to kernels and interconnect.

## What they're looking for

- 5+ years of professional system engineering experience
- Deep, hands-on experience with vLLM, SGLang (or a comparable serving engine) in production or at serious benchmark scale
- Strong software engineering skills in Python . You're comfortable working in large, performance-critical codebases
- A solid understanding of what drives LLM inference performance: batching, memory, parallelism, and the trade-offs between latency and throughput
- Experience with modern inference optimization techniques such as quantization, speculative decoding, or distributed serving
- Rigor in benchmarking and performance analysis, plus comfort with GPU profiling tools

Tags: Engineering
