Fast, reliable, reproducible AI with GPU live migration. Backed by Y Combinator.
About the role
Engineer solutions at client sites: Lead customer integrations. Install, configure, and deploy Cedana into SLURM, Kubernetes, and Dynamo environments. Drive product innovation from the field: Identify technical gaps while embedded with clients, then provide product feedback for new capabilities that become core product features.
What they're looking for
- Team management experience. Requires strong project and time management skills, delivering milestones on time, and effective
- 3-10 years of software engineering experience with a track record of configuring and managing SLURM deployments
- A multi-month enterprise or research deployment you led end-to-end, from scoping through signoff. You write effective status updates to keep your team updated and on schedule
- Production experience in standing up SLURM in a customer or research environment. You've configured slurmctld, slurmdbd, accounting, cgroup integration, and GPU resource selection
- Strong Linux fundamentals of systemd, cgroups v2, namespaces, networking, filesystems, firewalls, kernel module loading, PAM session modules. You can read strace and dmesg output and form a hypothesis
- Experience with Kubernetes operations including operators, CRDs, CNIs, device plugins, and node-level debugging. You've debugged a controller in production even if you haven't written one from scratch
More about this role
AI and HPC infrastructure suffers from scarcity and high costs, so when failures happen they are costly in terms of time and money. Cluster productivity directly determines research output and revenue. Achieving high utilization and throughput is increasingly challenging due to the complexity of workloads, hardware, and operations.
Cedana maximizes AI+HPC cluster utilization and reliability with automated GPU checkpointing infrastructure. We enable transparent and fast migration of GPU workloads across instances, without losing work. Workloads automatically migrate to achieve new levels of reliability and throughput while accelerating time to results. Our system is at the kernel/OS level, requiring no code or config changes, and works seamlessly with Kubernetes, SLURM, and NVIDIA Dynamo. Today, we're deploying into leading inference platforms, neoclouds, enterprise, and research clusters.
Cedana's founding team has spent over a decade making computation run fast, productively, and reliably for AI. Our research appears in NeurIPS and CVPR. We published some of the earliest formal methods for guaranteeing convergence in distributed training. At Shopify we've deployed warehouse...
Browse similar: AI jobs · AI startup jobs · Startup jobs · Remote jobs