Startups · AI

Software Engineer, GPU Kernels

River AI · Palo Alto, CA · On-site

← All jobs
About River AI

Develop frontier language models and agents. Training, reinforcement learning, and inference in one system, on River Cloud or your own GPU cluster. Backed by General Catalyst.

About the role

We are looking for exceptional GPU kernel engineers to build the compute primitives behind River’s training and inference infrastructure. Your goal is to make large models faster to train and more efficient to serve.

What they're looking for

  • Bachelor’s degree in Computer Science, Computer Engineering, or equivalent practical experience
  • Experience optimizing GPU kernels with CUDA, Triton, CUTLASS, CuTe, or comparable tools
  • Strong understanding of GPU architecture, memory hierarchies, and parallel execution
  • Proficiency in C++ and Python
  • Strong foundations in linear algebra, floating-point arithmetic, and numerical computing
  • Strong debugging and profiling skills, with a collaborative approach to engineering
More about this role

At River AI, our mission is to create personal AI owned and shaped by each individual. To achieve this, we are rewriting the entire stack from scratch: personal hardware for local inference, bespoke training infrastructure, next-generation UIs, and frontier deep learning research.

We are scientists, engineers, and builders from the industry's top tech companies and AI labs. We bring a proven track record of scaling consumer systems for hundreds of millions of users and architecting the pre-training infrastructure behind today's frontier models.

We are looking for exceptional GPU kernel engineers to build the compute primitives behind River’s training and inference infrastructure. Your goal is to make large models faster to train and more efficient to serve.

You will own performance-critical operations, including attention, matrix multiplication, mixture-of-experts execution, and low-precision computation. Working closely with researchers and systems engineers, you will identify bottlenecks, implement kernels, validate correctness, and bring improvements into production.

  • Build fast GPU kernels for attention, matrix multiplication, expert routing, and related operations.

-...

Read the full posting on River AI's site ↗

River API

Build your edge while you search

Free tools for founders and investors, plus VC Unfiltered, our take on startups, venture and the people who build them.