# Staff Software Engineer, Reliability at Metropolis

- Company: Metropolis
- What the company does: Metropolis transforms the parking experience with a computer vision platform that enables checkout-free payment.
- Company website: https://www.metropolis.io
- Type: Startups
- Level: Senior
- Location: Seattle, Washington, United States
- Work setup: On-site
- Pay: $204K to $249K base salary per year (USD)
- Posted: 2026-09-16
- Apply by: 2026-10-31
- Apply: https://job-boards.greenhouse.io/metropolis/jobs/7992483003
- Page: https://www.1752.vc/careers/jobs/metropolis-staff-software-engineer-reliability/

## About the role

Metropolis is seeking a Staff Software Engineer focused on Reliability to own reliability across the entire Metropolis platform and drive the comprehensive practices that ensure system availability, resilience, and observability for our mission-critical mobility infrastructure. In this role, you will build reliability from first principles, architecting failover systems, implementing chaos engineering, and improving our observability foundation to maintain 99.9%+ uptime as we scale to new markets.

## What they're looking for

- 8+ years of engineering experience including software engineering, reliability engineering, SRE practices, or production operations at scale
- Demonstrate expert-level reliability engineering skills including hands-on experience with multi-region architectures, failover automation, circuit breakers, chaos engineering, and disaster recovery
- Utilize production observability expertise with deep experience implementing monitoring, alerting, tracing, and logging systems at scale – specifically Datadog or similar APM platforms in high-load environments
- Apply strong systems thinking with proven ability to design resilient distributed systems that gracefully handle failures, network partitions, and external dependency outages
- Demonstrate database and data systems knowledge including replication strategies, backup/restore procedures, connection pooling, query optimization, and experience with both relational and NoSQL databases
- Leverage cloud platform expertise with production experience operating and ensuring reliability of systems on AWS including multi-region deployments, load balancing, and DNS-based failover

Tags: Application Development
