Startups

Head of Data Center Operations (Facility Operations)

Crux · Palo Alto · Remote

← All jobs
About Crux

Backed by a16z.

About the role

You will own critical facility operations for Crux’s entire fleet — every megawatt of electrical, mechanical, and controls infrastructure our TPU clusters depend on, across owned campuses and third-party colocation sites. You are accountable for uptime: the maintenance programs, operating procedures, compliance regime, and 24/7 site teams that keep high-density, liquid-cooled environments continuously available for AI workloads that do not tolerate thermal or power excursions.

What they're looking for

  • 15+ years in critical facility or data center operations, including 7+ years leading multi-site facility operations organizations at portfolio scale — you have owned uptime for a fleet
  • Built or scaled a facility operations program through rapid growth: standards libraries, CMMS deployments, commissioning-to-operations handoffs, and vendor/colo O&M management under contractual SLAs
  • Hyperscaler facility operations leadership (Meta, Google, AWS, Microsoft) or neocloud fleet operations during a rapid ramp
  • Direct-to-chip liquid cooling plant operations at production scale (CDUs, TCS/FWS loops, water chemistry programs)
  • Experience holding colocation operators to SLA across a leased portfolio, including audit and compliance program design
  • Commissioning leadership (Levels 1–5, IST) on mission-critical projects and first-year plant tuning
More about this role

Crux AI is a newly formed, U.S.-based integrated AI infrastructure company created to remove the physical and operational constraints on consequential AI ambitions. Crux brings together power, high-density data centers, TPU silicon, networking, orchestration software, and ongoing operations as one integrated system.

Crux is being capitalized to plan every layer together, develop each one to demanding standards, and operate the whole system with efficiency and reliability. That gives hyperscalers, frontier AI labs, sovereign customers, enterprises, and AI-native companies greater freedom to pursue the AI they are here to create.

You will own critical facility operations for Crux’s entire fleet — every megawatt of electrical, mechanical, and controls infrastructure our TPU clusters depend on, across owned campuses and third-party colocation sites. You are accountable for uptime: the maintenance programs, operating procedures, compliance regime, and 24/7 site teams that keep high-density, liquid-cooled environments continuously available for AI workloads that do not tolerate thermal or power excursions. You will build this organization from zero — campus facility managers, chief...

Read the full posting on Crux's site ↗

Development

Build your edge while you search

Free tools for founders and investors, plus VC Unfiltered, our take on startups, venture and the people who build them.