Startups · AI

Research Engineer - AI Performance & Kernel Optimization

Zyphra · San Francisco · On-site

← All jobs
About Zyphra

The future of intelligence is open.

About the role

As a Research Engineer - AI Performance & Kernel Optimization , you will improve and optimize the performance of our large-scale language model training and inference stacks. You will work closely with our pretraining and inference teams to identify bottlenecks, design and implement highly optimized kernels, and push the limits of throughput, latency, and hardware utilization across a range of accelerator platforms.

What they're looking for

  • Strong engineering aptitude for building reliable, high-performance systems
  • Excellent low-level performance intuition and the ability to reason about hardware-software interactions
  • Are excited to rapidly learn new systems, tools, and hardware environments
  • Excellent communication and collaboration skills, with the ability to work effectively across research and engineering teams
  • Enjoy diving deep into the weeds and hunting down the last 10–20% of performance
  • Experience writing highly performant GPU kernels at any level of abstraction–PTX, CUDA, HIP, Triton, or other kernel DSLs
More about this role

As a Research Engineer - AI Performance & Kernel Optimization , you will improve and optimize the performance of our large-scale language model training and inference stacks. You will work closely with our pretraining and inference teams to identify bottlenecks, design and implement highly optimized kernels, and push the limits of throughput, latency, and hardware utilization across a range of accelerator platforms. This role is suited for someone who enjoys deep systems work, cares about performance at every level of the stack, and is excited to translate low-level optimizations into meaningful gains for frontier-scale AI systems.

Kernel development and optimization for large-scale ML workloads, using any level of the stack from PTX/assembly to CUDA, HIP, Triton, or other GPU DSLs

Performance tuning for training and inference stacks across GPUs and other accelerators

Profiling and eliminating bottlenecks in memory movement, communication, scheduling, and compute utilization

Optimizing distributed training and inference systems for large MoE models, including large-scale model parallelism

Portability and optimization across non-NVIDIA hardware, with special interest in AMD...

Read the full posting on Zyphra's site ↗

R&D - Engineering

Build your edge while you search

Free tools for founders and investors, plus VC Unfiltered, our take on startups, venture and the people who build them.