The future of intelligence is open.
About the role
As a Research Engineer - AI Performance & Kernel Optimization , you will improve and optimize the performance of our large-scale language model training and inference stacks. You will work closely with our pretraining and inference teams to identify bottlenecks, design and implement highly optimized kernels, and push the limits of throughput, latency, and hardware utilization across a range of accelerator platforms.
What they're looking for
- Strong engineering aptitude for building reliable, high-performance systems
- Excellent low-level performance intuition and the ability to reason about hardware-software interactions
- Are excited to rapidly learn new systems, tools, and hardware environments
- Excellent communication and collaboration skills, with the ability to work effectively across research and engineering teams
- Enjoy diving deep into the weeds and hunting down the last 10–20% of performance
- Experience writing highly performant GPU kernels at any level of abstraction–PTX, CUDA, HIP, Triton, or other kernel DSLs
More about this role
As a Research Engineer - AI Performance & Kernel Optimization , you will improve and optimize the performance of our large-scale language model training and inference stacks. You will work closely with our pretraining and inference teams to identify bottlenecks, design and implement highly optimized kernels, and push the limits of throughput, latency, and hardware utilization across a range of accelerator platforms. This role is suited for someone who enjoys deep systems work, cares about performance at every level of the stack, and is excited to translate low-level optimizations into meaningful gains for frontier-scale AI systems.
Kernel development and optimization for large-scale ML workloads, using any level of the stack from PTX/assembly to CUDA, HIP, Triton, or other GPU DSLs
Performance tuning for training and inference stacks across GPUs and other accelerators
Profiling and eliminating bottlenecks in memory movement, communication, scheduling, and compute utilization
Optimizing distributed training and inference systems for large MoE models, including large-scale model parallelism
Portability and optimization across non-NVIDIA hardware, with special interest in AMD...
Browse similar: AI jobs · AI startup jobs · Startup jobs · San Francisco Bay Area