Startups · AI

LLM Inference Engineer

Hippocratic AI · Menlo Park, CA · On-site

← All jobs
About Hippocratic AI

Hippocratic AI builds the safest generative AI healthcare agent for health systems, payors, and pharma. Over 180 million clinical interactions across 1,000+ use cases with 60+ partners worldwide. Backed by General Catalyst, Kleiner Perkins and a16z.

About the role

Design and implement multi-node serving architectures for distributed LLM inference Apply advanced quantization techniques (FP4/FP6) to reduce model footprint while preserving quality

What they're looking for

  • Experience optimizing LLM inference systems at scale
  • Proven expertise with distributed serving architectures for large language models
  • Hands-on experience implementing quantization techniques for transformer models
  • Strong understanding of modern inference optimization methods, including:
  • Speculative decoding techniques with draft models
  • Eagle speculative decoding approaches
More about this role

As HAI's LLM Inference Engineer, you will own the serving infrastructure that determines whether our breakthrough healthcare AI reaches patients efficiently and reliably. You'll optimize the systems that translate raw model capability into sub-100ms responses—making the difference between conversational experiences that feel natural and those that feel broken. This role exists because inference optimization at scale is where research meets reality: your work directly determines latency, cost, and availability for millions of patient conversations across healthcare systems.

Own your first major outcome: By day 90, you will have shipped a measurable improvement to our inference serving stack (reduced latency, improved throughput, or optimized cost per inference), validated the gains across our production deployment scenarios, and established the performance optimization roadmap that will guide infrastructure investment.

Drive lasting impact: At 12 months, you will have designed and deployed advanced serving architectures (disaggregated inference, optimized caching, speculative decoding) that meaningfully improve patient experience and operational efficiency, contributed novel...

Read the full posting on Hippocratic AI's site ↗

Research & Development

Build your edge while you search

Free tools for founders and investors, plus VC Unfiltered, our take on startups, venture and the people who build them.