RadixArk builds large-scale inference and training systems for the entire AI community, making frontier-level AI infrastructure open and accessible. Backed by Accel.
About the role
RadixArk is seeking a Member of Technical Staff — Inference to push the limits of large-scale AI inference. You will work on the core systems that serve frontier models at scale, optimizing performance, latency, throughput, and cost across thousands of GPUs. This role sits at the intersection of systems engineering, ML infrastructure, and performance optimization.
What they're looking for
- 5+ years of experience in systems engineering, ML infrastructure, or performance-critical backend systems
- Strong expertise in large-scale inference systems for LLMs or generative models
- Deep understanding of GPU architecture and performance characteristics
- Experience optimizing latency- and throughput-critical production systems
- Strong knowledge of distributed systems and networking fundamentals
- Proficiency in Python, Rust, C++, or Go for production systems
More about this role
RadixArk is seeking a Member of Technical Staff — Inference to push the limits of large-scale AI inference.
You will work on the core systems that serve frontier models at scale, optimizing performance, latency, throughput, and cost across thousands of GPUs. This role sits at the intersection of systems engineering, ML infrastructure, and performance optimization.
Your work will directly shape how state-of-the-art models are deployed and experienced by users worldwide.
This is a deeply technical, high-impact role for engineers who enjoy working close to the hardware–software boundary and solving performance-critical problems at scale.
5+ years of experience in systems engineering, ML infrastructure, or performance-critical backend systems
Strong expertise in large-scale inference systems for LLMs or generative models
Deep understanding of GPU architecture and performance characteristics
Experience optimizing latency- and throughput-critical production systems
Strong knowledge of distributed systems and networking fundamentals
Proficiency in Python, Rust, C++, or Go for production systems
Experience profiling and optimizing compute-intensive workloads
Strong debugging skills across...
Browse similar: AI jobs · AI startup jobs · Startup jobs · San Francisco Bay Area