RadixArk builds large-scale inference and training systems for the entire AI community, making frontier-level AI infrastructure open and accessible. Backed by Accel.
About the role
As a Technical Program Manager at RadixArk, you'll drive the execution of complex, cross-functional programs across our inference and training infrastructure. You'll partner closely with Product Management, Research, and Engineering to turn ambitious technical roadmaps into shipped reality, coordinating across kernel teams, distributed systems engineers, and external partners to deliver infrastructure that serves billions of tokens daily and coordinates 10,000+ GPU training runs.
What they're looking for
- Bachelor's or Master's degree in Computer Science, Engineering, or a related technical field
- 4+ years of direct experience in Technical Program Management, Engineering Management, or a senior engineering role with significant program ownership, in a software or infrastructure company
- Strong technical fluency in systems software, distributed systems, or AI/ML infrastructure, able to read code, follow architecture discussions, and challenge technical assumptions productively
- Demonstrated track record shipping complex, multi-team programs on time, including managing dependencies, risks, and scope changes
- Excellent written and verbal communication skills, able to drive alignment across engineers, executives, and external partners
- Direct experience shipping AI/ML infrastructure such as inference engines, training frameworks, GPU kernels, distributed schedulers, or model serving platforms
More about this role
As a Technical Program Manager at RadixArk, you'll drive the execution of complex, cross-functional programs across our inference and training infrastructure. You'll partner closely with Product Management, Research, and Engineering to turn ambitious technical roadmaps into shipped reality, coordinating across kernel teams, distributed systems engineers, and external partners to deliver infrastructure that serves billions of tokens daily and coordinates 10,000+ GPU training runs.
This role is for someone who thrives at the intersection of deep technical understanding and rigorous program execution. You'll own the "how" and "when" of our most critical initiatives.
Drive end-to-end execution of large-scale, cross-functional programs spanning inference engines (e.g., SGLang), training frameworks (e.g., Miles), and hardware integration efforts.
Define program structure, including milestones, dependencies, critical paths, risks, and success criteria. Maintain a clear source of truth for status across all stakeholders.
Run design reviews, sprint planning, release readiness reviews, and post-mortems. Ensure decisions are documented and follow-ups are closed out.
Identify and unblock...
Browse similar: AI jobs · AI startup jobs · Startup jobs · San Francisco Bay Area