Startups · AI

Member of Technical Staff, Kernels

Inception Labs · San Mateo, United States · On-site

← All jobs
About Inception Labs

We are leveraging diffusion technology to develop a new generation of LLMs. Our dLLMs are much faster and more efficient than traditional autoregressive LLMs. Backed by AI Grant and Amplify.

About the role

We're looking for engineers and scientists to design, optimize, and maintain the compute foundations that power large-scale language model training and inference. You will develop high-performance ML kernels, enable efficient low-precision arithmetic, and improve the distributed compute stack that makes training and serving large models possible.

What they're looking for

  • Design and implement custom ML kernels (CUDA, CuTe, Triton) for core dLLM operations such as attention, matrix multiplication, gating, and normalization, optimized for modern GPU architectures
  • Design compute primitives to reduce memory bandwidth bottlenecks and improve kernel efficiency
  • Contribute to infrastructure stability and scalability, ensuring reproducibility, consistency across precision formats, and high utilization of compute resources
  • BS/MS/PhD in Computer Science, Engineering, or a related field (or equivalent experience)
  • Proficiency in CUDA, CuTe, Triton, or other GPU programming frameworks
  • Understanding of ML frameworks (PyTorch, TensorFlow) from a systems perspective
More about this role

Inception creates the world’s fastest, most efficient AI models. Our Mercury model is the world’s fastest reasoning LLM and first commercially available diffusion LLM, delivering 5x greater speed and efficiency than today’s LLMs, with best-in-class quality.

We are the AI researchers and engineers behind such breakthrough AI technologies as diffusion models, flash attention, and DPO.

The Role

We're looking for engineers and scientists to design, optimize, and maintain the compute foundations that power large-scale language model training and inference. You will develop high-performance ML kernels, enable efficient low-precision arithmetic, and improve the distributed compute stack that makes training and serving large models possible.

Key Responsibilities

  • Design and implement custom ML kernels (CUDA, CuTe, Triton) for core dLLM operations such as attention, matrix multiplication, gating, and normalization, optimized for modern GPU architectures.
  • Design compute primitives to reduce memory bandwidth bottlenecks and improve kernel efficiency.
  • Contribute to infrastructure stability and scalability, ensuring reproducibility, consistency across precision formats, and high...

Read the full posting on Inception Labs's site ↗

Build your edge while you search

Free tools for founders and investors, plus VC Unfiltered, our take on startups, venture and the people who build them.