About the role
We are seeking a Senior Site Reliability Engineer to own the reliability, performance, and resilience of the cloud infrastructure powering Garner’s products and AI/ML workloads. This role sits on our Platform Engineering team. You will run the machine: defining and upholding SLOs, leading incident response, and driving the automation and standards that let every Garner engineer ship faster and more reliably.
What they're looking for
- 4+ years of hands-on experience operating production cloud infrastructure at scale in an SRE, DevOps, or platform engineering role
- Deep expertise with Kubernetes and Terraform in a cloud-first environment (AWS preferred)
- A strong track record with production observability: defining SLOs, building monitoring and alerting, and leading incident response and blameless post-incident reviews
- Strong software engineering fundamentals in Python or Go, applied to infrastructure automation (experience with Kubernetes APIs a plus)
- Experience driving cloud cost-efficiency and performance optimization across compute, storage, and networking
- Experience supporting AI/ML or data-intensive workloads in production is a plus
More about this role
Garner is on a mission to transform the U.S. healthcare system — and we’re the only proven player doing exactly that. We partner with employers to redesign how healthcare works: applying 550+ proprietary clinical metrics across 80+ specialties to a dataset of 320M+ patients to identify the best-performing doctors, then using compelling incentives to steer members to the care that helps them get healthier, faster.
The result is a rare “win win” — better care and lower costs for both members and employers. In just five years, our work has helped over 2.5 million people access higher-quality care and saved $1B in healthcare costs. We recently raised our Series E and have doubled five years running. If you've ever wanted your work to solve a problem that touches every person in this country, this is the opportunity to do exactly that. You'd be joining a team fundamentally reimagining healthcare in the U.S. — and using AI to scale that impact further and faster than anyone else can.
We are seeking a Senior Site Reliability Engineer to own the reliability, performance, and resilience of the cloud infrastructure powering Garner’s products and AI/ML workloads. This role sits on our...
Browse similar: Startup jobs · Remote jobs