RadixArk builds large-scale inference and training systems for the entire AI community, making frontier-level AI infrastructure open and accessible. Backed by Accel.
About the role
RadixArk is looking for a Member of Technical Staff Cluster Infrastructure to architect and scale the core compute platform that powers frontier-level AI training and inference. You will design and operate highly reliable, high-performance GPU/TPU clusters, build next-generation scheduling and resource management systems, and push the limits of large-scale distributed infrastructure for AI workloads.
What they're looking for
- 5+ years of experience in distributed systems, infrastructure, or large-scale compute platforms
- Strong background in distributed systems design and systems architecture
- Deep experience with cluster management systems (Kubernetes, Slurm, Ray, or custom schedulers)
- Hands-on experience with GPU/TPU infrastructure in production environments
- Strong Linux systems and networking fundamentals
- Proficiency in Go, Rust, C++, or Python for production systems
More about this role
RadixArk is looking for a Member of Technical Staff Cluster Infrastructure to architect and scale the core compute platform that powers frontier-level AI training and inference.
You will design and operate highly reliable, high-performance GPU/TPU clusters, build next-generation scheduling and resource management systems, and push the limits of large-scale distributed infrastructure for AI workloads.
This role focuses on deep systems engineering across cluster architecture, networking, scheduling, and performance optimization. Your work will directly impact how efficiently frontier AI models are trained and served.
5+ years of experience in distributed systems, infrastructure, or large-scale compute platforms
Strong background in distributed systems design and systems architecture
Deep experience with cluster management systems (Kubernetes, Slurm, Ray, or custom schedulers)
Hands-on experience with GPU/TPU infrastructure in production environments
Strong Linux systems and networking fundamentals
Proficiency in Go, Rust, C++, or Python for production systems
Experience debugging complex multi-layer issues across hardware, OS, networking, and distributed services
Proven ability to...
Browse similar: AI jobs · AI startup jobs · Startup jobs · San Francisco Bay Area