Train and scale AI on NVIDIA VR200 NVL 72, GB300 NVL 72, B300, B200, H200, H100, and and more GPUs. Launch on-demand instances or reserve a cluster. Get started.
About the role
Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers to enterprises and hyperscalers. Lambda's mission is to make compute as ubiquitous as electricity and give everyone the power of superintelligence. One person, one GPU. Build and operate monitoring and alerting for cluster health — fabric, GPU, power/thermal, and job-level signals — to detect and respond to issues proactively
More about this role
Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers to enterprises and hyperscalers. Lambda's mission is to make compute as ubiquitous as electricity and give everyone the power of superintelligence. One person, one GPU.
If you'd like to build the world's best AI cloud, join us.
What You’ll Do
Build and operate monitoring and alerting for cluster health — fabric, GPU, power/thermal, and job-level signals — to detect and respond to issues proactively
Remotely deploy and configure large-scale HPC clusters for AI workloads using automation wherever possible
Automate cluster lifecycle: operating systems, firmware, drivers, and networking, managed as code (Ansible, Terraform) rather than by hand
Create runbooks and automated remediations for common cluster failure modes, designed so Support and HPC Support can run them safely
Troubleshoot and resolve cluster issues across InfiniBand/RoCE, NCCL, GPU-direct, fabric, switching, and power — working closely with on-site deployment teams
Participate in on-call rotations and lead incident response for cluster-level problems
Contribute to...
Browse similar: AI jobs · AI startup jobs · Startup jobs · Remote jobs · San Francisco Bay Area