# Member of Technical Staff, Local Inference & Kernels at Sonder

- Company: Sonder
- What the company does: Backed by Greylock and a16z speedrun.
- Company website: https://www.sonder.com/
- Type: Startups (AI role)
- Level: Senior
- Location: New York City
- Work setup: On-site
- Pay: $250K to $300K base salary per year (USD)
- Posted: 2026-09-18
- Apply by: 2026-11-02
- Apply: https://jobs.ashbyhq.com/sonder/ba8b22ab-eddd-496b-b661-631426de2ddf
- Page: https://www.1752.vc/careers/jobs/sonder-member-of-technical-staff-local-inference-and-kernels/

## About the role

We are looking for an inference and performance engineer to own the systems layer between our models and the hardware they run on. You will make our models faster, smaller, and more power-efficient, working from model execution and quantization down to memory layouts and custom GPU kernels. We are starting with Apple Silicon and macOS, using MLX and Metal.

## What they're looking for

- You do not need prior Apple Silicon experience. We care about demonstrated ability to understand hardware and make neural networks run substantially better on it
- - Deep experience in ML inference, GPU programming, or numerical computing, with concrete examples of improvements you have shipped
- - Strong C++ skills and hands-on experience writing kernels in Metal, CUDA, Triton, or a comparable accelerator programming environment
- - A working understanding of GPU architecture: memory hierarchies, bandwidth, SIMD execution, register pressure, occupancy, and synchronization. You can explain how these affect a kernel's performance
- - Experience optimizing matrix multiplication, attention, or similarly demanding operations, including validating numerical correctness across shapes and precision formats
- - An understanding of quantization and mixed precision, and the ability to measure their effects on memory, execution time, and model quality

Tags: Technical Staff
