Startups · AI

Kernel Engineer (Custom Silicon), Hardware

River AI · Palo Alto, CA; Austin, TX · On-site

← All jobs
About River AI

Develop frontier language models and agents. Training, reinforcement learning, and inference in one system, on River Cloud or your own GPU cluster. Backed by General Catalyst.

About the role

We are looking for exceptional performance and kernel generation engineers to build the foundational compute engine for our high-performance custom silicon. In this role, you will design and implement robust kernel generators that programmatically emit optimized low-level assembly code for our greenfield hardware architecture.

What they're looking for

  • Bachelor’s degree in Computer Engineering, Computer Science, Electrical Engineering, or a related field, and 5+ years of practical industry experience in low-level performance programming
  • Deep understanding of hardware programming models (e.g., CUDA, Triton, CUTLASS, or custom accelerator assembly) and a proven track record of shipping highly optimized kernels
  • Advanced knowledge of Computer Architecture, including vector units, execution pipelines, register files, and complex memory hierarchies (caches, SRAM, HBM/DRAM)
  • Proficiency in modern C++ for building robust, scalable meta-programming and code-generation frameworks
  • Strong mathematical foundation in linear algebra operations and deep learning primitives
  • A highly collaborative mindset to push boundaries and co-design effectively with hardware and compiler teams
More about this role

At River, our mission is to create personal AI owned and shaped by each individual. To achieve this, we are rewriting the entire stack from scratch: personal hardware for local inference, custom training infrastructure, next-generation UIs, and frontier deep learning research.

We are scientists, engineers, and builders from the industry's top tech companies and AI labs. We bring a proven track record of scaling consumer systems for hundreds of millions of users and architecting the pre-training infrastructure behind today's frontier models.

We are looking for exceptional performance and kernel generation engineers to build the foundational compute engine for our high-performance custom silicon. In this role, you will design and implement robust kernel generators that programmatically emit optimized low-level assembly code for our greenfield hardware architecture.

You will bridge the gap between high-level compilation and raw hardware capability, pushing our custom architecture to its absolute theoretical limits for critical deep learning operations (including GEMMs, FlashAttention, and custom activations). You will collaborate closely up and down the stack with compiler engineers,...

Read the full posting on River AI's site ↗

Personal AI Hardware

Build your edge while you search

Free tools for founders and investors, plus VC Unfiltered, our take on startups, venture and the people who build them.