Specific Intelligence for Your Business. Backed by Kleiner Perkins and Lux.
About the role
As a research scientist, you will design, implement, and optimize the large-scale training infrastructure that powers our frontier reinforcement learning stack. This is systems work at the edge of what's possible, training state-of-the-art models for our enterprise partners. Frontier systems are exciting but brittle, and require both performance and correctness to train models effectively. You'll work closely with researchers to make our RL stack reliable, fast, and capable of running for days without intervention.
What they're looking for
- Experience programming with and managing training jobs on large-scale GPU systems
- Fearlessness and curiosity to understand all levels of the training stack
- Bias toward fast implementation, paired with a high bar for reliability and efficiency
- Familiarity with open-weights models (architecture and inference)
- Background in reinforcement learning or integration of inference with RL training loops
More about this role
As a research scientist, you will design, implement, and optimize the large-scale training infrastructure that powers our frontier reinforcement learning stack. This is systems work at the edge of what's possible, training state-of-the-art models for our enterprise partners. Frontier systems are exciting but brittle, and require both performance and correctness to train models effectively. You'll work closely with researchers to make our RL stack reliable, fast, and capable of running for days without intervention.
Design and optimize our RL training and inference pipelines across large GPU clusters
Build tooling and observability that lets researchers and customers inspect, profile, and debug training runs
Implement systems with an eye toward how they affect ML (low precision numerics, distributed training edge cases, etc.)
Partner with researchers to bring frontier post-training capabilities into production deployments
Experience programming with and managing training jobs on large-scale GPU systems
Fearlessness and curiosity to understand all levels of the training stack
Bias toward fast implementation, paired with a high bar for reliability and efficiency
Familiarity with...
Browse similar: AI jobs · AI startup jobs · Startup jobs · San Francisco Bay Area