Develop frontier language models and agents. Training, reinforcement learning, and inference in one system, on River Cloud or your own GPU cluster. Backed by General Catalyst.
About the role
We are looking for exceptional systems engineers to build the high-performance engines that train our models. Your goal is to make training at River fast, reliable, and massively scalable. You will take ownership of our core infrastructure stack; from writing custom GPU kernels to managing clusters of thousands of nodes, ensuring our researchers can focus on science rather than system bottlenecks.
What they're looking for
- Bachelor’s degree in Computer Science, Computer Engineering, or equivalent practical industry experience
- Deep expertise in systems-level languages (C, C++, or Rust) with a track record of writing performant, maintainable code
- Strong foundation in computer architecture, memory management, and concurrent programming
- Exceptional debugging skills, especially when tackling complex, non-deterministic issues in distributed environments
- A highly collaborative mindset and a bias for action to push boundaries across the stack
- Hands-on experience with modern AI frameworks (e.g., PyTorch, JAX) and tooling for large-scale model training
More about this role
At River AI, our mission is to create personal AI owned and shaped by each individual. To achieve this, we are rewriting the entire stack from scratch: personal hardware for local inference, bespoke training infrastructure, next-generation UIs, and frontier deep learning research.
We are scientists, engineers, and builders from the industry's top tech companies and AI labs. We bring a proven track record of scaling consumer systems for hundreds of millions of users and architecting the pre-training infrastructure behind today's frontier models.
We are looking for exceptional systems engineers to build the high-performance engines that train our models. Your goal is to make training at River fast, reliable, and massively scalable.
You will take ownership of our core infrastructure stack; from writing custom GPU kernels to managing clusters of thousands of nodes, ensuring our researchers can focus on science rather than system bottlenecks.
- Architect and deploy fault-tolerant distributed systems for training and inference workloads across clusters with thousands of nodes.
- Design high-performance kernels to maximize tensor operation efficiency, memory throughput, and networking...
Browse similar: AI jobs · AI startup jobs · Startup jobs · San Francisco Bay Area