Develop frontier language models and agents. Training, reinforcement learning, and inference in one system, on River Cloud or your own GPU cluster. Backed by General Catalyst.
About the role
We are looking for exceptional inference systems engineers to build the engines that serve large models through the River API. Your goal is to deliver fast, reliable inference while making efficient use of GPU compute and memory. You will take ownership of the serving runtime, from request scheduling and continuous batching to KV-cache management, distributed model execution, and checkpoint loading. Your work will support both customer-facing inference and the sampling workloads that power reinforcement learning.
What they're looking for
- Bachelor’s degree in Computer Science, Computer Engineering, or equivalent practical experience
- Experience building inference engines or performance-sensitive distributed services
- Strong understanding of transformer inference, GPU memory, concurrency, and networking
- Proficiency in Python and C++ or Rust
- Strong debugging and profiling skills across models, runtimes, and services
- A collaborative mindset and strong ownership of engineering outcomes
More about this role
At River AI, our mission is to create personal AI owned and shaped by each individual. To achieve this, we are rewriting the entire stack from scratch: personal hardware for local inference, bespoke training infrastructure, next-generation UIs, and frontier deep learning research.
We are scientists, engineers, and builders from the industry's top tech companies and AI labs. We bring a proven track record of scaling consumer systems for hundreds of millions of users and architecting the pre-training infrastructure behind today's frontier models.
We are looking for exceptional inference systems engineers to build the engines that serve large models through the River API. Your goal is to deliver fast, reliable inference while making efficient use of GPU compute and memory.
You will take ownership of the serving runtime, from request scheduling and continuous batching to KV-cache management, distributed model execution, and checkpoint loading. Your work will support both customer-facing inference and the sampling workloads that power reinforcement learning.
Working closely with GPU kernel engineers, researchers, and infrastructure engineers, you will bring new models into production...
Browse similar: AI jobs · AI startup jobs · Startup jobs · San Francisco Bay Area