Anthropic is an AI safety and research company that's working to build reliable, interpretable, and steerable AI systems. Backed by Accel, Bessemer and General Catalyst.
About the role
As a Capacity Deployment Lead within Anthropic's Data Center Operations (DCO) organization, you will support and scale how we bring new compute capacity online across a growing fleet of partner-operated data centers. Deployment runs from construction kickoff to ready for service (RFS). You are engaged during planning and construction, and on-site and hands-on from EFA through infrastructure delivery and receiving, integration QA, power-on and network turn-up, burn-in and validation, and acceptance.
What they're looking for
- Experience with GPU/accelerator or high-density liquid-cooled infrastructure, including bring-up of liquid-cooled racks
- Experience standing up operations at a new site or data hall, from commissioning handoff through first deployment
- Experience deploying more than one accelerator platform, or introducing a new hardware platform into an existing deployment program
- Experience delivering deployment outcomes inside partner-operated or colocation sites where on-the-floor operations are staffed via third parties
- Familiarity with dense accelerator network cabling, optics, and high-speed interconnect at rack and pod scale
- The annual compensation range for this role is listed below
More about this role
Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems.
As a Capacity Deployment Lead within Anthropic's Data Center Operations (DCO) organization, you will support and scale how we bring new compute capacity online across a growing fleet of partner-operated data centers. Deployment runs from construction kickoff to ready for service (RFS). You are engaged during planning and construction, and on-site and hands-on from EFA through infrastructure delivery and receiving, integration QA, power-on and network turn-up, burn-in and validation, and acceptance. You work with hardware vendors and operations partners for completion of each CU (capacity unit) with a structured handover into production operations and the maintenance program. The scope is compute: accelerator systems, racks, and the infrastructure around them - at the sites in service today and the ones coming online behind them.
This is a program and execution...
Browse similar: AI jobs · AI startup jobs · Startup jobs · Remote jobs