Make intelligence open and accessible to all. Backed by Battery, Lightspeed and Sequoia.
About the role
Reflection's Compute Platform team keeps our compute layer healthy and highly available. We run a Kubernetes-based platform distributed across multiple neo-clouds, where multi-cloud scheduling, node health, and performance debugging at scale present genuinely hard systems problems.
What they're looking for
- Experience building, mentoring, and growing systems or infrastructure teams while staying technically hands-on. (Comfortable growing into leading a team of ~10 quickly if you haven't managed at that scale before.)
- Deep systems-level engineering experience with a focus on cluster-wide behavior and maintenance
- Strong coding ability and the credibility to earn the technical trust of a strong team
- Depth in at least one of orchestration, storage, or GPU hardware — with the ability to learn the rest. Deep GPU knowledge beyond standard Kubernetes (e.g., NCCL) is a plus, not a prerequisite
- Alignment with a Kubernetes-first architecture
- Cloud storage expertise — managing high-performance data products (like VAST) across multiple data centers and handling datasets and checkpointing at scale — is a plus
More about this role
Reflection is a research lab making intelligence open and accessible for everyone to use, customize, and build on. We build open models that let anyone control their intelligence and help shape the future of AI. Our mission: make intelligence open and accessible to all.
Reflection's Compute Platform team keeps our compute layer healthy and highly available. We run a Kubernetes-based platform distributed across multiple neo-clouds, where multi-cloud scheduling, node health, and performance debugging at scale present genuinely hard systems problems.
As Compute Platform Lead, you'll provide front-line leadership of the team that builds and operates this layer. You'll build, mentor, and grow a team of strong systems engineers, guide the technical and architectural decisions across multi-cloud scheduling, cluster management, and next-generation GPU deployments, and work closely with our training teams to co-design fault tolerance, node health checks, and remediation. You'll stay close enough to the systems to make targeted contributions as an individual contributor and to maintain a deep understanding of the compute fleet our largest training runs depend on. Managing vendors — and the...
Browse similar: AI jobs · AI startup jobs · Startup jobs · San Francisco Bay Area