Build, test, and deploy advanced voice AI agents in minutes with Vapi. The platform for developers creating conversational voice AI. Backed by Bessemer, Peak XV and Y Combinator.
About the role
30 Day : Join the oncall rotation. Walk the 15 stability-gap incidents and turn the patterns into a prioritized reliability backlog. Define the first set of SLOs for the call-completion path. 60 Day : Stand up error budgets and SLO-based alerting in Chronosphere/Prometheus for the highest-impact services. Run the first proper load test against provider rate limits and per-org concurrency. Tune autoscaling for wscaler / workerpool-cron-scaler.
What they're looking for
- Must-haves
- You’ve run incident command and postmortem discipline at scale on a real oncall rotation
- You’ve operated SLOs and error budgets in Chronosphere, Prometheus, Grafana, or Datadog
- You’ve done capacity planning and load testing for production systems with real users
- You’re fluent in Kubernetes production ops: pod crash diagnosis, HPA/VPA tuning, PodDisruptionBudgets, graceful shutdown
- You know backpressure and autoscaling patterns — KEDA, custom metrics scaling
More about this role
Voice AI that resolves, not transfers
Powering 1 billion calls for companies like Amazon Ring, Intuit, ServiceTitan, and New York Life
Trusted by 1 million developers building the future of voice agents
Backed by Peak XV, Bessemer, Kleiner Perkins, M12, Y Combinator, and more with $72M raised
Try talking to Vapi now!
99.99% call completion is the number this role drives. Vapi runs live phone calls — a p99 spike means callers drop. We’ve had 15 stability-gap outages worth learning from, and we need someone who runs incident command, owns SLOs and error budgets, and builds the reliability culture from scratch.
This is not a bash-and-YAML role. You’ll ship code (Go or TypeScript) for services that monitor and manage the platform: auto-remediation, capacity forecasters, oncall tooling. Capacity planning, load testing, and KEDA-based autoscaling for Vapi’s wscaler and workerpool-cron-scaler are on your plate.
30 Day : Join the oncall rotation. Walk the 15 stability-gap incidents and turn the patterns into a prioritized reliability backlog. Define the first set of SLOs for the call-completion path.
60 Day : Stand up error budgets and SLO-based alerting in Chronosphere/Prometheus for the...
Browse similar: AI jobs · AI startup jobs · Startup jobs · Remote jobs · San Francisco Bay Area