Startups · AI

Head of Engineering

Inferact · San Francisco · On-site

← All jobs
About Inferact

Inferact is a startup founded by creators and core maintainers of vLLM, the most popular open-source LLM inference engine. Our mission is to grow vLLM as the world. Backed by Sequoia and Redpoint.

About the role

Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware, a position that took years to build.

What they're looking for

  • Bachelor's degree or equivalent experience in computer science, engineering, machine learning, systems, or a related field
  • Engineering leadership experience building and scaling highly specialized teams in LLM inference, ML systems, GPU or accelerator software, distributed systems, or closely related infrastructure
  • Deep technical credibility at the inference layer, including hands-on understanding of inference runtimes, GPU or accelerator optimization, kernels, memory and communication bottlenecks, and hardware-software tradeoffs
  • Ability to distinguish core inference-engine work from the routing, orchestration, and application layers above it, with opinions grounded in direct technical experience
  • A strong record of recruiting, assessing, and retaining senior engineers, staff-level ICs, PhDs, and research-adjacent engineers in a production engineering environment
  • Experience translating technically ambitious work into clear priorities, accountable ownership, execution plans, and durable engineering operating mechanisms
More about this role

Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware, a position that took years to build.

We're looking for a Head of Engineering to build and lead the organization developing the systems that power vLLM and Inferact. This role requires an engineering leader with genuine technical credibility at the inference layer—someone who understands GPU and accelerator performance, inference runtimes, ML systems optimization, and hardware-software co-design deeply enough to earn the trust of exceptional staff-level engineers.

You'll partner closely with the founders to scale a senior-heavy, highly specialized engineering team while preserving the technical rigor, speed, and ownership that made vLLM successful. You'll recruit and develop rare ML systems talent, translate ambitious research and infrastructure work into a focused execution plan, strengthen how teams operate, and help Inferact deliver reliable, high-performance inference across models, hardware, and deployment environments.

Bachelor's degree...

Read the full posting on Inferact's site ↗

Research & Engineering

Build your edge while you search

Free tools for founders and investors, plus VC Unfiltered, our take on startups, venture and the people who build them.