Startups · AI

Member of Technical Staff, Forward Deployed Engineer

Inferact · San Francisco · On-site

← All jobs
About Inferact

Inferact is a startup founded by creators and core maintainers of vLLM, the most popular open-source LLM inference engine. Our mission is to grow vLLM as the world. Backed by Sequoia and Redpoint.

About the role

Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware, a position that took years to build. Found on 1752vc Careers, the job board for startup and VC roles.

What they're looking for

  • Bachelor's degree or equivalent experience in computer science, engineering, systems, machine learning, or similar
  • Strong software engineering ability in Python, Go, TypeScript, or similar, with experience building production-quality integrations, tooling, services, automation, or prototypes
  • Hands-on experience deploying or operating ML systems, model serving, AI infrastructure, cloud platforms, Kubernetes, or high-scale backend systems in production
  • Ability to work directly with sophisticated customer engineering teams, understand ambiguous technical environments, and personally drive implementations and debugging to resolution
  • Strong systems debugging skills across application, runtime, infrastructure, networking, identity, storage, observability, and distributed-system boundaries
  • Ability to reason about latency, throughput, batching, model/runtime compatibility, scaling, reliability, and cost tradeoffs in production inference environments
More about this role

Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware, a position that took years to build.

We're looking for a Forward Deployed Engineer to make Inferact successful inside real customer environments. You'll work directly with customers to deploy, integrate, debug, and optimize vLLM-powered inference systems across cloud, Kubernetes, GPU, networking, and model-serving environments.

This is a hands-on engineering role, not a traditional pre-sales position. You'll move from architecture discussions to implementation, own difficult production problems end-to-end, and work closely with core product and engineering teams to turn what you learn in the field into reusable product capabilities. Your work will directly affect customer time-to-value and how Inferact's platform evolves.

Bachelor's degree or equivalent experience in computer science, engineering, systems, machine learning, or similar.

Strong software engineering ability in Python, Go, TypeScript, or similar, with experience building...

Read the full posting on Inferact's site ↗

Research & Engineering

Build your edge while you search

Free tools for founders and investors, plus VC Unfiltered, our take on startups, venture and the people who build them.