AI infrastructure with on-demand GPUs and serverless compute. Run training, inference, and batch workloads on the cloud with Runpod. Backed by AI Grant.
About the role
Own Core Infrastructure & SRE: Lead multiple engineering teams responsible for Site Reliability Engineering, networking, and storage. Establish rigorous SRE practices, driving SLA/SLO definitions, incident response, observability, and automated remediation.
What they're looking for
- Engineering Leadership Experience: 7+ years leading software, infrastructure, SRE, or networking teams, including managing managers and multiple squads, with a proven record of scaling high-availability cloud environments
- Deep Infrastructure Expertise: 8+ years building and operating large-scale distributed systems, bare-metal infrastructure, or public/private cloud platforms
- HPC & Advanced Networking: Proven hands-on background or strong architectural understanding of ultra-low latency networking. Deep familiarity with InfiniBand and/or RoCE , spine-leaf architectures, and global WAN routing protocols (BGP)
- Storage Systems Knowledge: Experience building, operating, or tuning high-performance distributed storage systems and parallel file systems (e.g., Ceph, Lustre, Weka, NVMe-oF) capable of handling heavy AI/ML I/O loads
- SRE / DevOps Culture: Strong foundation in reliability engineering, infrastructure-as-code (Terraform, Ansible), container orchestration (Kubernetes), and modern observability stacks
- Remote-First Operating Excellence: Experience building culture, accountability, and momentum across distributed technical teams
More about this role
Runpod is the AI Developer Cloud. More than one million developers, from indie researchers to teams running frontier models in production, use Runpod to experiment, train, fine-tune, deploy, and scale AI on one platform. The platform has processed more than 20 billion inference requests. We closed a $100M Series A in June 2026. We're at an inflection point for AI infrastructure, and we're building the platform the next generation of developers will depend on. We're a small, remote-first team. We take ownership seriously, move fast, and ship work that more than a million developers rely on every day. We're looking for people who care deeply, build with urgency, and want to matter at scale.
Learn more in our CEO's funding announcement: https://www.runpod.io/blog/one-million-developers .
We’re looking for a Director of Infrastructure Engineering to lead and scale Runpod’s core cloud and bare-metal environments. This role owns the critical foundational layers of our platform—Site Reliability Engineering (SRE), global networking, High-Performance Computing (HPC) networks, and distributed storage engines. You’ll build the operating rhythm, culture, and technical direction that ensures...
Browse similar: AI jobs · AI startup jobs · Startup jobs · Remote jobs