Startups · AI

Member of Technical Staff — Inference-TPU

RadixArk · Palo Alto, CA · On-site

← All jobs
About RadixArk

RadixArk builds large-scale inference and training systems for the entire AI community, making frontier-level AI infrastructure open and accessible. Backed by Accel.

About the role

RadixArk is looking for a Member of Technical Staff — TPU Systems to build high-performance inference and training systems using JAX, XLA, and Pallas. You'll push model workloads to their limits on TPU hardware, working on SGLang-JAX and other critical infrastructure that enables efficient deployment of frontier models on Google's tensor processing units.

What they're looking for

  • 3+ years experience building production ML systems utilizing JAX/Torch, XLA, or TPU-focused frameworks
  • Bachelor's or Master's degree in Computer Science, Electrical Engineering, or equivalent industry experience
  • Deep understanding of XLA internals preferred: HLO, MLIR, operator fusion, SPMD partitioning, and sharding strategies
  • Strong performance tuning instincts across compiler and runtime layers
  • Experience with distributed inference systems (e.g. SGLang, vLLM) or training frameworks (e.g. Miles, Alpa, Pathways)
  • Proficiency in Python with demonstrated ability to write high-performance, production-quality code
More about this role

RadixArk is looking for a Member of Technical Staff — TPU Systems to build high-performance inference and training systems using JAX, XLA, and Pallas. You'll push model workloads to their limits on TPU hardware, working on SGLang-JAX and other critical infrastructure that enables efficient deployment of frontier models on Google's tensor processing units.

  • 3+ years experience building production ML systems utilizing JAX/Torch, XLA, or TPU-focused frameworks.
  • Bachelor's or Master's degree in Computer Science, Electrical Engineering, or equivalent industry experience
  • Deep understanding of XLA internals preferred: HLO, MLIR, operator fusion, SPMD partitioning, and sharding strategies.
  • Strong performance tuning instincts across compiler and runtime layers
  • Experience with distributed inference systems (e.g. SGLang, vLLM) or training frameworks (e.g. Miles, Alpa, Pathways)
  • Proficiency in Python with demonstrated ability to write high-performance, production-quality code
  • Experience writing custom GPU/TPU/AI Accelerator kernels. Familiarity with Pallas for kernel development is strongly preferred.
  • Build high-performance inference and training systems using JAX/XLA/Pallas,...

Read the full posting on RadixArk's site ↗

Member of Technical Staff

Build your edge while you search

Free tools for founders and investors, plus VC Unfiltered, our take on startups, venture and the people who build them.