# Member of Technical Staff, Site Reliablity Engineer at Vapi

- Company: Vapi
- What the company does: Build, test, and deploy advanced voice AI agents in minutes with Vapi. The platform for developers creating conversational voice AI. Backed by Bessemer, Peak XV and Y Combinator.
- Company website: https://vapi.ai/
- Type: Startups (AI role)
- Level: Senior
- Location: San Francisco
- Work setup: Remote
- Pay: $200K to $270K base salary per year (USD)
- Posted: 2026-06-03
- Apply by: 2026-10-08
- Apply: https://jobs.ashbyhq.com/vapi/4b6abd59-ac74-40e9-8cab-74253078aaf4
- Page: https://www.1752.vc/careers/jobs/vapi-member-of-technical-staff-site-reliablity-engineer/

## About the role

30 Day : Join the oncall rotation. Walk the 15 stability-gap incidents and turn the patterns into a prioritized reliability backlog. Define the first set of SLOs for the call-completion path. 60 Day : Stand up error budgets and SLO-based alerting in Chronosphere/Prometheus for the highest-impact services. Run the first proper load test against provider rate limits and per-org concurrency. Tune autoscaling for wscaler / workerpool-cron-scaler.

## What they're looking for

- Must-haves
- You’ve run incident command and postmortem discipline at scale on a real oncall rotation
- You’ve operated SLOs and error budgets in Chronosphere, Prometheus, Grafana, or Datadog
- You’ve done capacity planning and load testing for production systems with real users
- You’re fluent in Kubernetes production ops: pod crash diagnosis, HPA/VPA tuning, PodDisruptionBudgets, graceful shutdown
- You know backpressure and autoscaling patterns — KEDA, custom metrics scaling

Tags: Engineering
