Startups · AI

Member of Technical Staff — Inference-Core Engine

RadixArk · Palo Alto, CA · On-site

← All jobs
About RadixArk

RadixArk builds large-scale inference and training systems for the entire AI community, making frontier-level AI infrastructure open and accessible. Backed by Accel.

About the role

RadixArk is seeking a Member of Technical Staff — Inference to push the limits of large-scale AI inference. You will work on the core systems that serve frontier models at scale, optimizing performance, latency, throughput, and cost across thousands of GPUs. This role sits at the intersection of systems engineering, ML infrastructure, and performance optimization.

What they're looking for

  • 5+ years of experience in systems engineering, ML infrastructure, or performance-critical backend systems
  • Strong expertise in large-scale inference systems for LLMs or generative models
  • Deep understanding of GPU architecture and performance characteristics
  • Experience optimizing latency- and throughput-critical production systems
  • Strong knowledge of distributed systems and networking fundamentals
  • Proficiency in Python, Rust, C++, or Go for production systems
More about this role

RadixArk is seeking a Member of Technical Staff — Inference to push the limits of large-scale AI inference.

You will work on the core systems that serve frontier models at scale, optimizing performance, latency, throughput, and cost across thousands of GPUs. This role sits at the intersection of systems engineering, ML infrastructure, and performance optimization.

Your work will directly shape how state-of-the-art models are deployed and experienced by users worldwide.

This is a deeply technical, high-impact role for engineers who enjoy working close to the hardware–software boundary and solving performance-critical problems at scale.

5+ years of experience in systems engineering, ML infrastructure, or performance-critical backend systems

Strong expertise in large-scale inference systems for LLMs or generative models

Deep understanding of GPU architecture and performance characteristics

Experience optimizing latency- and throughput-critical production systems

Strong knowledge of distributed systems and networking fundamentals

Proficiency in Python, Rust, C++, or Go for production systems

Experience profiling and optimizing compute-intensive workloads

Strong debugging skills across...

Read the full posting on RadixArk's site ↗

Member of Technical Staff

Build your edge while you search

Free tools for founders and investors, plus VC Unfiltered, our take on startups, venture and the people who build them.