# Site Reliability Engineer at Runpod

- Company: Runpod
- What the company does: AI infrastructure with on-demand GPUs and serverless compute. Run training, inference, and batch workloads on the cloud with Runpod. Backed by AI Grant.
- Company website: https://www.runpod.io/
- Type: Startups (AI role)
- Level: Mid level
- Location: Remote - USA
- Work setup: Remote
- Pay: $150K to $200K base salary per year (USD)
- Posted: 2026-07-10
- Apply by: 2026-10-08
- Apply: https://jobs.ashbyhq.com/runpod/1d14340c-c9bd-4754-80f4-5c83980cd413
- Page: https://www.1752.vc/careers/jobs/runpod-site-reliability-engineer/

## About the role

Increase platform uptime and reduce incident frequency and duration Improve MTTR through better tooling, automation, and runbooks

## What they're looking for

- 5+ years of experience in SRE, Reliability Engineering, or Production Engineering
- Strong Linux systems and Networking expertise
- Experience managing containerized production systems
- Strong understanding of distributed systems and failure modes
- Experience defining and managing SLIs/SLOs
- Proven incident response and postmortem leadership experience

Tags: Engineering
