RadixArk builds large-scale inference and training systems for the entire AI community, making frontier-level AI infrastructure open and accessible. Backed by Accel.
About the role
As a Member of Technical Staff, Training, you will design, build, and operate the distributed systems behind large-scale model post-training — spanning training, inference, and orchestration, with a focus on the performance, correctness, scalability, and reliability of workloads running across large GPU clusters.
What they're looking for
- 3+ years of experience building or operating distributed machine learning systems, large-scale training infrastructure, or high-performance inference systems
- Hands-on experience with post-training systems, training backends, or inference systems for large language models (e.g., Megatron-LM, FSDP, SGLang, TensorRT-LLM, vLLM)
- Experience in at least two of the following areas:
- Numerical correctness or low precision
- Stability, reliability, or fault tolerance
- Post-training algorithm recipes and orchestration infrastructure for large training runs
More about this role
As a Member of Technical Staff, Training, you will design, build, and operate the distributed systems behind large-scale model post-training — spanning training, inference, and orchestration, with a focus on the performance, correctness, scalability, and reliability of workloads running across large GPU clusters.
This role suits engineers who move fluidly across modeling recipes, complex infrastructure, and low-level systems, identify bottlenecks in distributed workloads, and translate experimental requirements into robust software.
- Design, build, and operate distributed training, rollout, and orchestration systems for large-scale LLM and multimodal post-training across multi-GPU, multi-node environments.
- Profile and optimize performance across the full-stack — model implementation, parallelism strategies, communication libraries, and GPU kernels — to improve throughput, latency, memory efficiency, hardware utilization, and cost.
- Investigate numerical correctness and low-precision issues in distributed training and inference, including train–inference consistency for reinforcement learning.
- Improve the reliability of long-running workloads through checkpointing, fault...
Browse similar: AI jobs · AI startup jobs · Startup jobs · San Francisco Bay Area