Power AI workloads with AMD Instinct™ GPUs. Scale models faster with high-performance, dedicated cloud compute.
About the role
We’re looking for a Software Engineer (Back-end) to join our platform team during an exciting phase of growth. In this role, you’ll own the end-to-end automation of provisioning, configuring, and operating large-scale GPU clusters across bare metal, Kubernetes, and Slurm environments.
What they're looking for
- 5+ years in infrastructure engineering or platform engineering
- Knowledge of GPU workload infrastructure
- Experience with RoCE networking automation
- Experience with GitOps tools such as ArgoCD
- Experience with CI/CD tools such as GitHub Actions and Argo Workflows
- Experience with Ansible and Terraform
More about this role
Our mission is simple: deliver seamless, secure, reliable, and resilient AI compute at scale. We've built a versatile cloud platform that eliminates infrastructure barriers, empowering builders to focus on innovation instead of fighting their stack. Because breakthrough AI should move at the speed of ideas, not infrastructure.
We’re looking for a Software Engineer (Back-end) to join our platform team during an exciting phase of growth. In this role, you’ll own the end-to-end automation of provisioning, configuring, and operating large-scale GPU clusters across bare metal, Kubernetes, and Slurm environments. This is a hands-on technical role focused on building the tooling and pipelines that bring hundreds of GPU nodes online reliably and repeatably—working closely with cross-functional partners to support business objectives while upholding our standards for excellence, collaboration, and impact.
Build and maintain fully automated pipelines for provisioning bare metal GPU clusters from zero to production
Automate Slurm and Kubernetes cluster lifecycle—bootstrapping, upgrades, node provisioning, and decommissioning at scale
Develop and maintain infrastructure for GPU node...
Browse similar: Startup jobs · Remote jobs