Backed by Greylock and a16z speedrun.
About the role
We are looking for an inference and performance engineer to own the systems layer between our models and the hardware they run on. You will make our models faster, smaller, and more power-efficient, working from model execution and quantization down to memory layouts and custom GPU kernels. We are starting with Apple Silicon and macOS, using MLX and Metal.
What they're looking for
- You do not need prior Apple Silicon experience. We care about demonstrated ability to understand hardware and make neural networks run substantially better on it
- - Deep experience in ML inference, GPU programming, or numerical computing, with concrete examples of improvements you have shipped
- - Strong C++ skills and hands-on experience writing kernels in Metal, CUDA, Triton, or a comparable accelerator programming environment
- - A working understanding of GPU architecture: memory hierarchies, bandwidth, SIMD execution, register pressure, occupancy, and synchronization. You can explain how these affect a kernel's performance
- - Experience optimizing matrix multiplication, attention, or similarly demanding operations, including validating numerical correctness across shapes and precision formats
- - An understanding of quantization and mixed precision, and the ability to measure their effects on memory, execution time, and model quality
More about this role
Sonder is an applied AI lab building models that learn how people work.
Today, AI largely depends on people explaining what they are doing and remembering when to ask for help. We are building private models that live on your computer, understand work as it happens, and learn the patterns, preferences, and judgment behind how you operate.
Our first product, Twin, brings this intelligence to the Mac. It builds a memory of your work and uses that context to offer help at the right moment.
We are looking for an inference and performance engineer to own the systems layer between our models and the hardware they run on.
You will make our models faster, smaller, and more power-efficient, working from model execution and quantization down to memory layouts and custom GPU kernels. We are starting with Apple Silicon and macOS, using MLX and Metal.
Our models run throughout the working day, sharing memory and compute with the applications someone is using. Sustained power draw, memory bandwidth, and responsiveness under contention matter as much as peak throughput. A faster kernel matters when it makes the whole system better.
You will work closely with researchers to co-design models and...
Browse similar: AI jobs · AI startup jobs · Startup jobs