Startups

Senior Infrastructure Engineer, SRE

Rocket Money · San Francisco, CA, Washington, D.C., New York City, NY, Remote (USA) · Remote

← All jobs
About Rocket Money

Rocket Money is a subscription manager that helps find and cancel unwanted subscriptions, and helps you create a custom budget to track monthly spending and expenses. Backed by Accel.

About the role

Additional information: Salary range of $ 150,000 - $185,000 /year + bonus + benefits. Base pay offered may vary depending on job-related knowledge, skills, and experience. Found on 1752vc Careers, the job board for startup and VC roles.

What they're looking for

  • You have 5+ years of hands-on cloud or infrastructure engineering experience, with substantial time spent on reliability and production operations at scale
  • You have defined SLIs and SLOs for real production services, and can talk about what changed as a result. What got fixed, what got deprioritized, and what you got wrong the first time
  • You have hands-on experience with an observability platform in production, Datadog strongly preferred
  • You're comfortable writing code (Python, Go, TypeScript, or similar) for internal tooling, production debugging, and automation
  • You write production Terraform and are comfortable in AWS, and when production breaks you can find the problem and fix it
  • You have built or operated a disaster recovery plan: you set the recovery goals, wrote the failover and restore steps, and ran the drills that proved it works
More about this role

Rocket Money’s mission is to empower people to live their best financial lives. Rocket Money offers members a unique understanding of their finances and a suite of valuable services that save them time and money – ultimately giving them a leg up on their financial journey.

We're looking to expand our Cloud Infrastructure team with a Senior Infrastructure Engineer, SRE to lead the reliability and operational evolution of our platform. We run hundreds of services in production, which enable us to process billions of transactions, consume multiple terabytes of data, and produce hundreds of millions of logs per day, and our reliability practice needs to evolve to match our growing scale. This includes:

  • Building and improving the reliability and resiliency of our systems and services
  • Establishing SLIs, SLOs, and error budgets for our most critical services and user journeys, and reviewing them regularly with the teams that own them
  • Owning and evolving our disaster recovery strategy: recovery objectives, failover and restore paths, and regular exercises that prove they work
  • Partnering with product engineering teams so they can own and operate their own services, with metrics...

Read the full posting on Rocket Money's site ↗

Engineering

Build your edge while you search

Free tools for founders and investors, plus VC Unfiltered, our take on startups, venture and the people who build them.