Backed by a16z.
About the role
We are recruiting founding Senior Site Reliability and Software Engineering Managers to build and lead our initial fleet reliability engineering teams in Palo Alto, CA. In this organization, there is no separate software development team. Your team owns the software, control plane, telemetry, and automated remediation controllers that keep multi-gigawatt TPU clusters provisioned, resilient, and continuously executing customer AI workloads.
What they're looking for
- 10+ years of software or infrastructure engineering experience, with 3+ years managing engineering teams owning direct production SLAs and on-call
- Deep SRE Discipline: Grounded in foundational SRE principles (SLOs, error budgets, blameless postmortems) paired with a strict "code over heroics" mindset
- Extensive AI/ML Adoption: Active utilization of AI agents and automated LLM/ML workflows in modern software engineering and diagnostic operations
- Senior Talent Magnet: Track record of attracting, evaluating, developing and leading unusually senior software engineers who thrive in fast-paced, high-stakes environments
- Hyperscaler / Neocloud Scale: SRE or fleet leadership at a hyperscaler (Google, AWS, Meta, MSFT) or neocloud (CoreWeave, Lambda, Nebius, Nscale) during rapid fleet ramps
- Accelerated Compute: Direct TPU experience or large-scale GPU cluster ops (NCCL collective debugging, RDMA/GPU-Direct, Slurm/Kubernetes AI schedulers)
More about this role
Crux AI is a newly formed, U.S.-based integrated AI infrastructure company created to remove the physical and operational constraints on consequential AI ambitions. Crux brings together power, high-density data centers, TPU silicon, networking, orchestration software, and ongoing operations as one integrated system.
Crux is being capitalized to plan every layer together, develop each one to demanding standards, and operate the whole system with efficiency and reliability. That gives hyperscalers, frontier AI labs, sovereign customers, enterprises, and AI-native companies greater freedom to pursue the AI they are here to create.
Crux AI is led by CEO, Ben Treynor Sloss , who spent over two decades in executive technical leadership at Google and founded the Site Reliability Engineering (SRE) disciplin e. At Crux AI, we treat operations fundamentally as a software engineering problem.
We are recruiting founding Senior Site Reliability and Software Engineering Managers to build and lead our initial fleet reliability engineering teams in Palo Alto, CA.
In this organization, there is no separate software development team. Your team owns the software, control plane, telemetry, and...
Browse similar: Startup jobs · Founding team roles · Remote jobs · San Francisco Bay Area