Startups · AI

Staff Site Reliability Engineer

Hippocratic AI · Menlo Park, CA · On-site

← All jobs
About Hippocratic AI

Hippocratic AI builds the safest generative AI healthcare agent for health systems, payors, and pharma. Over 180 million clinical interactions across 1,000+ use cases with 60+ partners worldwide. Backed by General Catalyst, Kleiner Perkins and a16z.

About the role

We're looking for a Staff Site Reliability Engineer who is equally at home writing production software and running the infrastructure it lives on — and who wants to take ownership of one of the hardest, highest-leverage problems on our platform: intelligently managing a large fleet of GPU-backed models.

What they're looking for

  • Must-Have
  • 10+ years of professional experience across site reliability / DevOps engineering and software engineering
  • Computer Science Degree Required from a top CS program
  • Strong software engineering fundamentals — you build orchestration and scheduling systems in Python and/or Go, not just configure off-the-shelf tools
  • Experience designing systems that make decisions from operational metrics — collecting signals, interpreting them, and driving control loops such as autoscaling, load shedding, or admission control
  • Deep experience with infrastructure automation and CI/CD (Terraform, GitLab CI/CD, or similar)
More about this role

We're looking for a Staff Site Reliability Engineer who is equally at home writing production software and running the infrastructure it lives on — and who wants to take ownership of one of the hardest, highest-leverage problems on our platform: intelligently managing a large fleet of GPU-backed models.

We run nearly 30 models across heterogeneous hardware, and keeping that fleet fast, reliable, and cost-effective is a serious engineering challenge. You'll build the GPU management and scheduling platform that sits at the center of it — collecting utilization and load metrics, interpreting what they actually mean, and using them to make real-time decisions about admission control and scaling. The goal: route and schedule inference calls so we use our capacity efficiently without exceeding it, and scale model replicas up and down automatically as demand shifts.

This is a senior role for someone with a decade in the field who can move fluidly between systems engineering and software development, and who is excited to own a complex, evolving system end to end.

Design and build our GPU management and scheduling platform — the system that decides when, where, and how inference calls run...

Read the full posting on Hippocratic AI's site ↗

Research & Development

Build your edge while you search

Free tools for founders and investors, plus VC Unfiltered, our take on startups, venture and the people who build them.