Train and scale AI on NVIDIA VR200 NVL 72, GB300 NVL 72, B300, B200, H200, H100, and and more GPUs. Launch on-demand instances or reserve a cluster. Get started.
About the role
As a Senior or Staff Software Engineer in Lambda’s Cloud Services Engineering organization, you will build and operate the distributed systems that power Lambda’s GPU cloud. Our teams own platform capabilities across compute control planes, managed Kubernetes, cloud APIs, identity and access, usage metering and billing, capacity and orchestration, reliability, and developer-facing infrastructure.
What they're looking for
- 7 or more years of professional software engineering experience, or equivalent evidence of impact building production systems
- Experience designing, building, and operating backend services, distributed systems, infrastructure, or platform capabilities at meaningful scale
- Practical understanding of system design, data models, APIs, failure modes, performance, and the tradeoffs required to run reliable software in production
- A track record of owning complex work through delivery and operation, including testing, staged rollout, monitoring, incident response, and root-cause improvement
- Proven track record of aligning cross functional partners and gaining consensus around decisions and tradeoffs
More about this role
Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers to enterprises and hyperscalers. Lambda's mission is to make compute as ubiquitous as electricity and give everyone the power of superintelligence. One person, one GPU.
If you'd like to build the world's best AI cloud, join us.
*Note: This position requires presence in our San Francisco/San Jose/Bellevue office location 4 days per week; Lambda’s designated work from home day is currently Tuesday.
As a Senior or Staff Software Engineer in Lambda’s Cloud Services Engineering organization, you will build and operate the distributed systems that power Lambda’s GPU cloud. Our teams own platform capabilities across compute control planes, managed Kubernetes, cloud APIs, identity and access, usage metering and billing, capacity and orchestration, reliability, and developer-facing infrastructure.
You will turn large-scale GPU infrastructure into reliable, secure, customer-facing cloud services by building APIs, workflows, stateful controllers, schedulers, and operational tooling. You will be full cycle engineer, owning systems through...
Browse similar: AI jobs · AI startup jobs · Startup jobs · Remote jobs · San Francisco Bay Area