RadixArk builds large-scale inference and training systems for the entire AI community, making frontier-level AI infrastructure open and accessible. Backed by Accel.
About the role
RadixArk is looking for a Member of Technical Staff — TPU Systems to build high-performance inference and training systems using JAX, XLA, and Pallas. You'll push model workloads to their limits on TPU hardware, working on SGLang-JAX and other critical infrastructure that enables efficient deployment of frontier models on Google's tensor processing units.
What they're looking for
- 3+ years experience building production ML systems utilizing JAX/Torch, XLA, or TPU-focused frameworks
- Bachelor's or Master's degree in Computer Science, Electrical Engineering, or equivalent industry experience
- Deep understanding of XLA internals preferred: HLO, MLIR, operator fusion, SPMD partitioning, and sharding strategies
- Strong performance tuning instincts across compiler and runtime layers
- Experience with distributed inference systems (e.g. SGLang, vLLM) or training frameworks (e.g. Miles, Alpa, Pathways)
- Proficiency in Python with demonstrated ability to write high-performance, production-quality code
More about this role
RadixArk is looking for a Member of Technical Staff — TPU Systems to build high-performance inference and training systems using JAX, XLA, and Pallas. You'll push model workloads to their limits on TPU hardware, working on SGLang-JAX and other critical infrastructure that enables efficient deployment of frontier models on Google's tensor processing units.
- 3+ years experience building production ML systems utilizing JAX/Torch, XLA, or TPU-focused frameworks.
- Bachelor's or Master's degree in Computer Science, Electrical Engineering, or equivalent industry experience
- Deep understanding of XLA internals preferred: HLO, MLIR, operator fusion, SPMD partitioning, and sharding strategies.
- Strong performance tuning instincts across compiler and runtime layers
- Experience with distributed inference systems (e.g. SGLang, vLLM) or training frameworks (e.g. Miles, Alpa, Pathways)
- Proficiency in Python with demonstrated ability to write high-performance, production-quality code
- Experience writing custom GPU/TPU/AI Accelerator kernels. Familiarity with Pallas for kernel development is strongly preferred.
- Build high-performance inference and training systems using JAX/XLA/Pallas,...
Browse similar: AI jobs · AI startup jobs · Startup jobs · San Francisco Bay Area