# Senior Manager, Cluster Engineering & Deployment at TensorWave

- Company: TensorWave
- What the company does: Power AI workloads with AMD Instinct™ GPUs. Scale models faster with high-performance, dedicated cloud compute.
- Company website: https://tensorwave.com
- Type: Startups
- Level: Senior
- Location: Remote
- Work setup: Remote
- Posted: 2026-09-15
- Apply by: 2026-10-30
- Apply: https://jobs.ashbyhq.com/tensorwave/eaa9cd8f-7405-492e-9515-94bc8121d048
- Page: https://www.1752.vc/careers/jobs/tensorwave-senior-manager-cluster-engineering-and-deployment/

## About the role

The Senior Manager, Cluster Engineering & Deployment owns and runs the machine that turns delivered racks into accepted clusters: network bring-up, fabric cabling verification against port maps, GPU node integration with the fabric, cluster-level validation and burn-in (including RCCL/collective performance), and the acceptance gate into production. This is one of the most schedule-critical roles in the pillar cluster revenue starts when this team says a cluster is ready.

## What they're looking for

- 10+ years across network deployment, cluster/HPC bring-up, or large-scale infrastructure delivery, including managing engineers in a field/deployment setting
- Hands-on fabric bring-up experience at scale (hundreds of switches / thousands of links per deployment)
- Strong operational rigor: building and enforcing playbooks, gates, metrics, and blameless defect loops
- Team leadership with schedule accountability across multiple concurrent builds or sites
- GPU cluster validation experience (NCCL/RCCL benchmarking)
- Automation skills (Python, Ansible) applied to deployment

Tags: Data Center
