# Director of Infrastructure Engineering at Runpod

- Company: Runpod
- What the company does: AI infrastructure with on-demand GPUs and serverless compute. Run training, inference, and batch workloads on the cloud with Runpod. Backed by AI Grant.
- Company website: https://www.runpod.io/
- Type: Startups (AI role)
- Level: Principal and up
- Location: Remote - USA
- Work setup: Remote
- Pay: $225K to $325K base salary per year (USD)
- Posted: 2026-07-09
- Apply by: 2026-10-08
- Apply: https://jobs.ashbyhq.com/runpod/33f23f31-4751-49d8-b6ce-8e26df7104a6
- Page: https://www.1752.vc/careers/jobs/runpod-director-of-infrastructure-engineering/

## About the role

Own Core Infrastructure & SRE: Lead multiple engineering teams responsible for Site Reliability Engineering, networking, and storage. Establish rigorous SRE practices, driving SLA/SLO definitions, incident response, observability, and automated remediation.

## What they're looking for

- Engineering Leadership Experience: 7+ years leading software, infrastructure, SRE, or networking teams, including managing managers and multiple squads, with a proven record of scaling high-availability cloud environments
- Deep Infrastructure Expertise: 8+ years building and operating large-scale distributed systems, bare-metal infrastructure, or public/private cloud platforms
- HPC & Advanced Networking: Proven hands-on background or strong architectural understanding of ultra-low latency networking. Deep familiarity with InfiniBand and/or RoCE , spine-leaf architectures, and global WAN routing protocols (BGP)
- Storage Systems Knowledge: Experience building, operating, or tuning high-performance distributed storage systems and parallel file systems (e.g., Ceph, Lustre, Weka, NVMe-oF) capable of handling heavy AI/ML I/O loads
- SRE / DevOps Culture: Strong foundation in reliability engineering, infrastructure-as-code (Terraform, Ansible), container orchestration (Kubernetes), and modern observability stacks
- Remote-First Operating Excellence: Experience building culture, accountability, and momentum across distributed technical teams

Tags: Engineering
