Startups · AI

Member of Technical Staff — Training

RadixArk · Palo Alto, CA · On-site

← All jobs
About RadixArk

RadixArk builds large-scale inference and training systems for the entire AI community, making frontier-level AI infrastructure open and accessible. Backed by Accel.

About the role

As a Member of Technical Staff, Training, you will design, build, and operate the distributed systems behind large-scale model post-training — spanning training, inference, and orchestration, with a focus on the performance, correctness, scalability, and reliability of workloads running across large GPU clusters.

What they're looking for

  • 3+ years of experience building or operating distributed machine learning systems, large-scale training infrastructure, or high-performance inference systems
  • Hands-on experience with post-training systems, training backends, or inference systems for large language models (e.g., Megatron-LM, FSDP, SGLang, TensorRT-LLM, vLLM)
  • Experience in at least two of the following areas:
  • Numerical correctness or low precision
  • Stability, reliability, or fault tolerance
  • Post-training algorithm recipes and orchestration infrastructure for large training runs
More about this role

As a Member of Technical Staff, Training, you will design, build, and operate the distributed systems behind large-scale model post-training — spanning training, inference, and orchestration, with a focus on the performance, correctness, scalability, and reliability of workloads running across large GPU clusters.

This role suits engineers who move fluidly across modeling recipes, complex infrastructure, and low-level systems, identify bottlenecks in distributed workloads, and translate experimental requirements into robust software.

  • Design, build, and operate distributed training, rollout, and orchestration systems for large-scale LLM and multimodal post-training across multi-GPU, multi-node environments.
  • Profile and optimize performance across the full-stack — model implementation, parallelism strategies, communication libraries, and GPU kernels — to improve throughput, latency, memory efficiency, hardware utilization, and cost.
  • Investigate numerical correctness and low-precision issues in distributed training and inference, including train–inference consistency for reinforcement learning.
  • Improve the reliability of long-running workloads through checkpointing, fault...

Read the full posting on RadixArk's site ↗

Member of Technical Staff

Build your edge while you search

Free tools for founders and investors, plus VC Unfiltered, our take on startups, venture and the people who build them.