Startups

Senior Manager, Cluster Engineering & Deployment

TensorWave · Remote · Remote

← All jobs
About TensorWave

Power AI workloads with AMD Instinct™ GPUs. Scale models faster with high-performance, dedicated cloud compute.

About the role

The Senior Manager, Cluster Engineering & Deployment owns and runs the machine that turns delivered racks into accepted clusters: network bring-up, fabric cabling verification against port maps, GPU node integration with the fabric, cluster-level validation and burn-in (including RCCL/collective performance), and the acceptance gate into production. This is one of the most schedule-critical roles in the pillar cluster revenue starts when this team says a cluster is ready.

What they're looking for

  • 10+ years across network deployment, cluster/HPC bring-up, or large-scale infrastructure delivery, including managing engineers in a field/deployment setting
  • Hands-on fabric bring-up experience at scale (hundreds of switches / thousands of links per deployment)
  • Strong operational rigor: building and enforcing playbooks, gates, metrics, and blameless defect loops
  • Team leadership with schedule accountability across multiple concurrent builds or sites
  • GPU cluster validation experience (NCCL/RCCL benchmarking)
  • Automation skills (Python, Ansible) applied to deployment
More about this role

Our mission is simple: deliver seamless, secure, reliable, and resilient AI compute at scale. We've built a versatile cloud platform that eliminates infrastructure barriers, empowering builders to focus on innovation instead of fighting their stack. Because breakthrough AI should move at the speed of ideas, not infrastructure.

The Senior Manager, Cluster Engineering & Deployment owns and runs the machine that turns delivered racks into accepted clusters: network bring-up, fabric cabling verification against port maps, GPU node integration with the fabric, cluster-level validation and burn-in (including RCCL/collective performance), and the acceptance gate into production. This is one of the most schedule-critical roles in the pillar cluster revenue starts when this team says a cluster is ready.

Own the cluster deployment playbook and drive its evolution: staged bring-up, automated config push, link/optics validation, cabling verification against L1 port maps, and fault triage during deployment windows.

Lead deployment engineering across concurrent cluster builds, through team leads and on-site engineers; coordinate daily with Data Center Integration field teams and cabling...

Read the full posting on TensorWave's site ↗

Data Center

Build your edge while you search

Free tools for founders and investors, plus VC Unfiltered, our take on startups, venture and the people who build them.