Startups · AI

Software Engineer, Distributed Training

River AI · Palo Alto, CA · On-site

← All jobs
About River AI

Develop frontier language models and agents. Training, reinforcement learning, and inference in one system, on River Cloud or your own GPU cluster. Backed by General Catalyst.

About the role

We are looking for exceptional systems engineers to build the distributed training engines behind the River API. Your goal is to make fine-tuning and reinforcement learning fast, numerically correct, and reliable across large GPU clusters.

What they're looking for

  • Bachelor’s degree in Computer Science, Computer Engineering, or equivalent practical industry experience
  • Hands-on experience building or substantially improving distributed model-training systems
  • Strong proficiency in Python and a modern deep-learning framework, such as PyTorch or JAX
  • Solid understanding of backpropagation, optimizers, mixed-precision training, and GPU memory management
  • Strong debugging skills across concurrent execution, collective communication, and distributed failure recovery
  • A highly collaborative mindset and a bias for action to push boundaries across the stack
More about this role

At River AI, our mission is to create personal AI owned and shaped by each individual. To achieve this, we are rewriting the entire stack from scratch: personal hardware for local inference, bespoke training infrastructure, next-generation UIs, and frontier deep learning research.

We are scientists, engineers, and builders from the industry's top tech companies and AI labs. We bring a proven track record of scaling consumer systems for hundreds of millions of users and architecting the pre-training infrastructure behind today's frontier models.

We are looking for exceptional systems engineers to build the distributed training engines behind the River API. Your goal is to make fine-tuning and reinforcement learning fast, numerically correct, and reliable across large GPU clusters.

You will own the execution of training workloads, including gradient computation, optimizer updates, rollout coordination, and checkpoint recovery. Working closely with researchers and inference engineers, you will bring new learning methods into production and improve how efficiently models use compute.

  • Build and optimize distributed training for large dense and mixture-of-experts models, including...

Read the full posting on River AI's site ↗

River API

Build your edge while you search

Free tools for founders and investors, plus VC Unfiltered, our take on startups, venture and the people who build them.