Backed by a16z.
About the role
You will own critical facility operations for Crux’s entire fleet — every megawatt of electrical, mechanical, and controls infrastructure our TPU clusters depend on, across owned campuses and third-party colocation sites. You are accountable for uptime: the maintenance programs, operating procedures, compliance regime, and 24/7 site teams that keep high-density, liquid-cooled environments continuously available for AI workloads that do not tolerate thermal or power excursions.
What they're looking for
- 15+ years in critical facility or data center operations, including 7+ years leading multi-site facility operations organizations at portfolio scale — you have owned uptime for a fleet
- Built or scaled a facility operations program through rapid growth: standards libraries, CMMS deployments, commissioning-to-operations handoffs, and vendor/colo O&M management under contractual SLAs
- Hyperscaler facility operations leadership (Meta, Google, AWS, Microsoft) or neocloud fleet operations during a rapid ramp
- Direct-to-chip liquid cooling plant operations at production scale (CDUs, TCS/FWS loops, water chemistry programs)
- Experience holding colocation operators to SLA across a leased portfolio, including audit and compliance program design
- Commissioning leadership (Levels 1–5, IST) on mission-critical projects and first-year plant tuning
More about this role
Crux AI is a newly formed, U.S.-based integrated AI infrastructure company created to remove the physical and operational constraints on consequential AI ambitions. Crux brings together power, high-density data centers, TPU silicon, networking, orchestration software, and ongoing operations as one integrated system.
Crux is being capitalized to plan every layer together, develop each one to demanding standards, and operate the whole system with efficiency and reliability. That gives hyperscalers, frontier AI labs, sovereign customers, enterprises, and AI-native companies greater freedom to pursue the AI they are here to create.
You will own critical facility operations for Crux’s entire fleet — every megawatt of electrical, mechanical, and controls infrastructure our TPU clusters depend on, across owned campuses and third-party colocation sites. You are accountable for uptime: the maintenance programs, operating procedures, compliance regime, and 24/7 site teams that keep high-density, liquid-cooled environments continuously available for AI workloads that do not tolerate thermal or power excursions. You will build this organization from zero — campus facility managers, chief...
Browse similar: Startup jobs · Remote jobs · San Francisco Bay Area