Startups · AI

Member of Technical Staff, Site Reliablity Engineer

Vapi · San Francisco · Remote

← All jobs
About Vapi

Build, test, and deploy advanced voice AI agents in minutes with Vapi. The platform for developers creating conversational voice AI. Backed by Bessemer, Peak XV and Y Combinator.

About the role

30 Day : Join the oncall rotation. Walk the 15 stability-gap incidents and turn the patterns into a prioritized reliability backlog. Define the first set of SLOs for the call-completion path. 60 Day : Stand up error budgets and SLO-based alerting in Chronosphere/Prometheus for the highest-impact services. Run the first proper load test against provider rate limits and per-org concurrency. Tune autoscaling for wscaler / workerpool-cron-scaler.

What they're looking for

  • Must-haves
  • You’ve run incident command and postmortem discipline at scale on a real oncall rotation
  • You’ve operated SLOs and error budgets in Chronosphere, Prometheus, Grafana, or Datadog
  • You’ve done capacity planning and load testing for production systems with real users
  • You’re fluent in Kubernetes production ops: pod crash diagnosis, HPA/VPA tuning, PodDisruptionBudgets, graceful shutdown
  • You know backpressure and autoscaling patterns — KEDA, custom metrics scaling
More about this role

Voice AI that resolves, not transfers

Powering 1 billion calls for companies like Amazon Ring, Intuit, ServiceTitan, and New York Life

Trusted by 1 million developers building the future of voice agents

Backed by Peak XV, Bessemer, Kleiner Perkins, M12, Y Combinator, and more with $72M raised

Try talking to Vapi now!

99.99% call completion is the number this role drives. Vapi runs live phone calls — a p99 spike means callers drop. We’ve had 15 stability-gap outages worth learning from, and we need someone who runs incident command, owns SLOs and error budgets, and builds the reliability culture from scratch.

This is not a bash-and-YAML role. You’ll ship code (Go or TypeScript) for services that monitor and manage the platform: auto-remediation, capacity forecasters, oncall tooling. Capacity planning, load testing, and KEDA-based autoscaling for Vapi’s wscaler and workerpool-cron-scaler are on your plate.

30 Day : Join the oncall rotation. Walk the 15 stability-gap incidents and turn the patterns into a prioritized reliability backlog. Define the first set of SLOs for the call-completion path.

60 Day : Stand up error budgets and SLO-based alerting in Chronosphere/Prometheus for the...

Read the full posting on Vapi's site ↗

Engineering

Build your edge while you search

Free tools for founders and investors, plus VC Unfiltered, our take on startups, venture and the people who build them.