Startups · AI

Member of Technical Staff — Reliability-CI Infrastructure

RadixArk · Palo Alto, CA · On-site

← All jobs
About RadixArk

RadixArk builds large-scale inference and training systems for the entire AI community, making frontier-level AI infrastructure open and accessible. Backed by Accel.

About the role

RadixArk is hiring a Member of Technical Staff — CI / Infrastructure to own the infrastructure that keeps SGLang moving. Our CI system runs 300+ GPU tests across NVIDIA, AMD, Intel, and Ascend hardware pools, gating every commit to one of the fastest-growing open-source LLM inference engines. When CI is green and fast, 100+ contributors ship with confidence. When it isn't, the entire project stalls. That bottleneck is your problem to solve.

What they're looking for

  • Deep Linux, Docker, GPU computing knowledge
  • Self-hosted runner management experience
  • Strong Bash and Python
  • Security mindset — CI supply chain risks, fork PR attack vectors, runner hardening
  • NVIDIA GPU drivers, CUDA, NCCL, InfiniBand/RDMA experience in CI contexts
  • Familiarity with ML inference workloads (model loading, KV cache, quantization)
More about this role

RadixArk is hiring a Member of Technical Staff — CI / Infrastructure to own the infrastructure that keeps SGLang moving. Our CI system runs 300+ GPU tests across NVIDIA, AMD, Intel, and Ascend hardware pools, gating every commit to one of the fastest-growing open-source LLM inference engines. When CI is green and fast, 100+ contributors ship with confidence. When it isn't, the entire project stalls. That bottleneck is your problem to solve.

You won't just maintain pipelines — you'll architect them. You'll replace brittle static thresholds with regression-based detection, harden runners against supply-chain attacks from fork PRs, and cut cycle times so contributors get feedback in minutes, not hours. You'll work directly with core maintainers, hardware partners, and the open-source community to keep the system that gates every merge request trustworthy, fast, and secure.

This is not a role for someone who wants to write CI YAML and walk away. It's for an engineer who treats CI infrastructure the way we treat serving infrastructure — as a system worth designing well.

  • Own CI reliability end-to-end — triage failures, distinguish real regressions from flaky tests and infra issues,...

Read the full posting on RadixArk's site ↗

Member of Technical Staff

Build your edge while you search

Free tools for founders and investors, plus VC Unfiltered, our take on startups, venture and the people who build them.