Startups · AI

Senior ML Systems Engineer, Inference

Runpod · Remote - USA · Remote

← All jobs
About Runpod

AI infrastructure with on-demand GPUs and serverless compute. Run training, inference, and batch workloads on the cloud with Runpod. Backed by AI Grant.

About the role

Define how we measure inference performance, including throughput, time to first token, inter-token latency, and cost per token, and build the tooling that makes those measurements rigorous and repeatable. Profile and diagnose performance problems across the serving stack, from scheduling and memory management down to kernels and interconnect.

What they're looking for

  • 5+ years of professional system engineering experience
  • Deep, hands-on experience with vLLM, SGLang (or a comparable serving engine) in production or at serious benchmark scale
  • Strong software engineering skills in Python . You're comfortable working in large, performance-critical codebases
  • A solid understanding of what drives LLM inference performance: batching, memory, parallelism, and the trade-offs between latency and throughput
  • Experience with modern inference optimization techniques such as quantization, speculative decoding, or distributed serving
  • Rigor in benchmarking and performance analysis, plus comfort with GPU profiling tools
More about this role

Runpod is the AI Developer Cloud. More than one million developers, from indie researchers to teams running frontier models in production, use Runpod to experiment, train, fine-tune, deploy, and scale AI on one platform. The platform has processed more than 20 billion inference requests. We closed a $100M Series A in June 2026. We're at an inflection point for AI infrastructure, and we're building the platform the next generation of developers will depend on.

We're a small, remote-first team. We take ownership seriously, move fast, and ship work that more than a million developers rely on every day. We're looking for people who care deeply, build with urgency, and want to matter at scale.

Learn more in our CEO's funding announcement: https://www.runpod.io/blog/one-million-developers .

We're looking for a ML Systems Engineer, Inference. We want Runpod to be the best place in the world to run LLM inference, meaning the fastest and the most cost-efficient. You'll lead that effort. You'll own LLM serving performance end to end. That means measuring it, understanding it, and improving it across models, hardware generations, and workloads. The work you ship will show up directly in the...

Read the full posting on Runpod's site ↗

Engineering

Build your edge while you search

Free tools for founders and investors, plus VC Unfiltered, our take on startups, venture and the people who build them.