# Member of Technical Staff, GPU Kernels at SF Tensor

- Company: SF Tensor
- What the company does: Train the models only you can build. SF Tensor provides one optimized, cross-vendor training stack for enterprise post-training and frontier pre-training. Backed by Y Combinator.
- Company website: https://sf-tensor.com
- Type: Startups (AI role)
- Level: Senior
- Location: San Francisco
- Work setup: On-site
- Pay: $285K to $315K base salary per year (USD)
- Posted: 2026-08-29
- Apply by: 2026-10-13
- Apply: https://jobs.ashbyhq.com/sf-tensor/2238bd4b-3fa6-44b7-bdac-e1e04ae2fc78
- Page: https://www.1752.vc/careers/jobs/sf-tensor-member-of-technical-staff-gpu-kernels/

## About the role

We build the fastest GPU compiler in the world. Most compilers have to preserve correctness at every transform, constraining how far they can search, while we prove correctness at the end instead, allowing us to search a far wider space, with agents, with RL, with anything that works and still guarantee the result. It's why we hold #1 on NVIDIA's own kernel benchmark across hundreds of production kernels.

## What they're looking for

- Someone with a track record of hand-writing kernels that match or beat vendor libraries
- Someone comfortable reading PTX, SASS, GCN/CDNA ISA or equivalent machine-level assembly
- Someone fluent with low-level profiling tools: Nsight Compute, Nsight Systems, rocprof, omniperf or their equivalent
- Someone with solid systems programming skills in C++ and CUDA or ROCm/HIP with a working understanding of how high-level ML operations map onto the hardware, including where the framework layers get in the way

Tags: Kernels
