# Principal Site Reliability Engineer at Tandem

- Company: Tandem
- What the company does: AI-native office leasing broker. Backed by Y Combinator and a16z.
- Company website: http://drivetandem.com
- Type: Startups
- Level: Principal and up
- Location: Remote US
- Work setup: Remote
- Pay: $165K to $185K base salary per year (USD)
- Posted: 2026-09-17
- Apply by: 2026-11-01
- Apply: https://tandemdiabetes.wd12.myworkdayjobs.com/en-US/tandemdiabetes/job/Remote-US/Principal-Site-Reliability-Engineer_JR101553
- Page: https://www.1752.vc/careers/jobs/tandem-principal-site-reliability-engineer/

## About the role

Informs capacity planning and scaling strategy with the Test team (who own load testing and modeling) and introduces proactive resilience testing such as game days once baseline support and observability are stable. Partners with application, architecture, and business stakeholders to align business continuity and disaster recovery capabilities with application requirements and recovery objectives.

## What they're looking for

- Demonstrated experience leading production support and incident management for production systems, including incident command during high-severity events
- Strong grounding in SRE principles: SLIs/SLOs, blameless postmortems, toil reduction, and treating reliability as an engineering discipline rather than a purely operational function
- Demonstrated experience owning on-call strategy, including rotation design, alert tuning, and escalation
- Expertise with Terraform or comparable IaC at scale: module design, state management, and policy-as-code guardrails
- Hands-on experience building CI/CD pipelines with reliability guardrails using tools such as GitHub Actions, Octopus Deploy, or Azure DevOps
- Deep experience with at least one major cloud platform (AWS, Azure, or GCP, [ preferred]) and with containerization and orchestration (Docker/Kubernetes)

