Fast, reliable, reproducible AI with GPU live migration. Backed by Y Combinator.
About the role
Validate and test automation: Our engineers contribute to and develop testing capabilities. This will be a core part of your initial work to establish your understanding of how our system works. Reliability and testing is a key part of our culture. Measure and optimize platform performance: Get Cedana to the theoretical maximum performance by understanding fundamental bottlenecks. Measure reliability, throughput and performance using our internal tools.
What they're looking for
- 5-10 years of software engineering experience with Linux Kernel
- CRIU + QEMU live migration + VFIO
- runc / containerd / OCI / namespaces / cgroups v2 / overlayfs
- Performance and latency engineering: characterizes throughput, jitter, and tail latency, uses tracing and profiling tooling (perf, ftrace, eBPF, or equivalent) to localize bottlenecks to a code path
- Systems-level QE: designs and owns automated test infrastructure, performance and regression harnesses, and reproduces kernel-level races and corner cases. Not manual or UI QA
- Enterprise Linux distribution environment (RHEL or equivalent): version matrices, backports, customer-grade triage
More about this role
AI and HPC infrastructure suffer from scarcity and high costs, so when failures occur, they are costly in time and money. Cluster productivity directly determines research output and revenue. Achieving high utilization and throughput is increasingly challenging due to the complexity of workloads, hardware, and operations.
Cedana maximizes AI+HPC cluster utilization and reliability with automated GPU checkpointing infrastructure. We enable transparent, fast migration of GPU workloads across instances without losing work. Workloads automatically migrate to achieve new levels of reliability and throughput while accelerating time to results. Our system is at the kernel/OS level, requiring no code or config changes, and works seamlessly with Kubernetes, SLURM, and NVIDIA Dynamo. Today, we're deploying into leading inference platforms, neoclouds, enterprise, and research clusters.
Cedana's founding team has spent over a decade making computation run fast, productively, and reliably for AI. Our research appears in NeurIPS and CVPR. We published some of the earliest formal methods for guaranteeing convergence in distributed training. At Shopify, we've developed a control plane for...
Browse similar: AI jobs · AI startup jobs · Startup jobs · Remote jobs