From bits to atoms. Backed by Accel, Lightspeed and a16z.
About the role
You’ll work alongside some of the world’s leading ML systems engineers, including leaders behind Megatron-LM, SGLang, Liger Kernel, TorchRec, CleanRL, TorchRL, and JAX-MD . We’re looking for exceptional ML Systems Engineers to build the agentic infrastructure powering our large-scale training, inference, and reinforcement learning. You’ll own critical pieces of the ML systems stack to maximize performance, scalability, reliability, and productivity for both engineers and AI agents.
What they're looking for
- Strong systems programming and performance engineering skills
- Experience building high-performance ML infrastructure at scale
- Ability to own complex technical problems end-to-end
- Strong coding ability and engineering judgment, including the ability to work effectively with AI agents to design, implement, test, and debug complex systems
- High ownership, fast execution, and a passion for pushing the frontier of AI systems and accelerating scientific discovery
More about this role
We're an AI and physical sciences company building state-of-the-art models to accelerate breakthroughs across materials, energy, and beyond. Backed by world-class investors and growing rapidly, we operate at the pace the frontier requires. Our team brings deep expertise, genuine ownership, and an insatiable drive to push the boundaries of what's scientifically possible.
You’ll work alongside some of the world’s leading ML systems engineers, including leaders behind Megatron-LM, SGLang, Liger Kernel, TorchRec, CleanRL, TorchRL, and JAX-MD .
We’re looking for exceptional ML Systems Engineers to build the agentic infrastructure powering our large-scale training, inference, and reinforcement learning. You’ll own critical pieces of the ML systems stack to maximize performance, scalability, reliability, and productivity for both engineers and AI agents.
Build and optimize large-scale training and reinforcement learning infrastructure while ensuring its correctness
Develop high-performance inference and serving systems
Design distributed runtimes and scheduling systems for complex ML workloads
Build secure and large-scale sandboxing and execution environments
Optimize memory, GPU kernels...
Browse similar: AI jobs · AI startup jobs · Startup jobs · San Francisco Bay Area