# Senior Site Reliability & Software Engineering Manager at Crux

- Company: Crux
- What the company does: Backed by a16z.
- Company website: https://www.cruxdata.com
- Type: Startups
- Level: Senior
- Location: Palo Alto
- Work setup: Remote
- Posted: 2026-09-15
- Apply by: 2026-10-30
- Apply: https://jobs.ashbyhq.com/crux/9c1c9e78-2c93-46c5-a3d6-62dc548e4914
- Page: https://www.1752.vc/careers/jobs/crux-senior-site-reliability-and-software-engineering-manager/

## About the role

We are recruiting founding Senior Site Reliability and Software Engineering Managers to build and lead our initial fleet reliability engineering teams in Palo Alto, CA. In this organization, there is no separate software development team. Your team owns the software, control plane, telemetry, and automated remediation controllers that keep multi-gigawatt TPU clusters provisioned, resilient, and continuously executing customer AI workloads.

## What they're looking for

- 10+ years of software or infrastructure engineering experience, with 3+ years managing engineering teams owning direct production SLAs and on-call
- Deep SRE Discipline: Grounded in foundational SRE principles (SLOs, error budgets, blameless postmortems) paired with a strict "code over heroics" mindset
- Extensive AI/ML Adoption: Active utilization of AI agents and automated LLM/ML workflows in modern software engineering and diagnostic operations
- Senior Talent Magnet: Track record of attracting, evaluating, developing and leading unusually senior software engineers who thrive in fast-paced, high-stakes environments
- Hyperscaler / Neocloud Scale: SRE or fleet leadership at a hyperscaler (Google, AWS, Meta, MSFT) or neocloud (CoreWeave, Lambda, Nebius, Nscale) during rapid fleet ramps
- Accelerated Compute: Direct TPU experience or large-scale GPU cluster ops (NCCL collective debugging, RDMA/GPU-Direct, Slurm/Kubernetes AI schedulers)

Tags: SRE
