# Senior Staff Data Center Operations Engineer, GPU Hardware Architecture at crusoe

- Company: crusoe
- What the company does: Crusoe provides next-gen AI infrastructure and cloud compute using an energy-first approach. Deploy AI workloads at scale with reliable performance and 24/7 support. Backed by Founders Fund, Felicis and Air Street.
- Company website: https://www.crusoe.ai
- Type: Startups (AI role)
- Level: Senior
- Location: San Francisco, CA - US
- Work setup: On-site
- Pay: $179K to $218K base salary per year (USD)
- Posted: 2026-09-23
- Apply by: 2026-11-07
- Apply: https://jobs.ashbyhq.com/crusoe/0a5354b7-a128-43a3-8f87-63a3109918b4
- Page: https://www.1752.vc/careers/jobs/crusoe-senior-staff-data-center-operations-engineer-gpu-hardware-architecture/

## About the role

Engineering Education & Design Support: Provide deep-dive technical guidance to the Data Center Engineering team on upcoming silicon (e.g., NVIDIA Blackwell/Rubin, AMD MI350/400). Ensure future facility designs for power, cooling, and rack-spacing are ready for 2000W+ per-chip densities.

## What they're looking for

- Silicon & Fabric Mastery: Expert-level knowledge of NVIDIA (Hopper/Blackwell/Rubin) and AMD (Instinct) architectures. Mastery of the physical and logical layers of NVLink, NVSwitch, and InfiniBand
- Infrastructure Bridge-Building: Ability to translate "Silicon Data Sheets" into "Mechanical Engineering Requirements." You can explain how a GPU's specific heat-load profile affects CDU sizing and secondary loop design
- Data-Driven Diagnostics: Proficient in Python, Go, or Bash to build telemetry and health-check tools (utilizing DCGM and ROCm ). Experience using large datasets or basic ML frameworks to build "Smart Monitoring" that filters critical health signals from noise
- Operational Reliability Analysis: Experience using failure telemetry to inform site-level sparing requirements and field-service workflows
- Thermal Management: Deep understanding of the operational realities of Direct-to-Chip (D2C) cooling, including fluid dynamics, pressure-drop curves, and the lifecycle of dripless couplings
- 10+ years in Hardware Engineering, Systems Architecture, or Data Center Infrastructure

Tags: Data Center Operations (DIG)
