We are leveraging diffusion technology to develop a new generation of LLMs. Our dLLMs are much faster and more efficient than traditional autoregressive LLMs. Backed by AI Grant and Amplify.
About the role
We're looking for engineers and scientists to design, optimize, and maintain the compute foundations that power large-scale language model training and inference. You will develop high-performance ML kernels, enable efficient low-precision arithmetic, and improve the distributed compute stack that makes training and serving large models possible.
What they're looking for
- Design and implement custom ML kernels (CUDA, CuTe, Triton) for core dLLM operations such as attention, matrix multiplication, gating, and normalization, optimized for modern GPU architectures
- Design compute primitives to reduce memory bandwidth bottlenecks and improve kernel efficiency
- Contribute to infrastructure stability and scalability, ensuring reproducibility, consistency across precision formats, and high utilization of compute resources
- BS/MS/PhD in Computer Science, Engineering, or a related field (or equivalent experience)
- Proficiency in CUDA, CuTe, Triton, or other GPU programming frameworks
- Understanding of ML frameworks (PyTorch, TensorFlow) from a systems perspective
More about this role
Inception creates the world’s fastest, most efficient AI models. Our Mercury model is the world’s fastest reasoning LLM and first commercially available diffusion LLM, delivering 5x greater speed and efficiency than today’s LLMs, with best-in-class quality.
We are the AI researchers and engineers behind such breakthrough AI technologies as diffusion models, flash attention, and DPO.
The Role
We're looking for engineers and scientists to design, optimize, and maintain the compute foundations that power large-scale language model training and inference. You will develop high-performance ML kernels, enable efficient low-precision arithmetic, and improve the distributed compute stack that makes training and serving large models possible.
Key Responsibilities
- Design and implement custom ML kernels (CUDA, CuTe, Triton) for core dLLM operations such as attention, matrix multiplication, gating, and normalization, optimized for modern GPU architectures.
- Design compute primitives to reduce memory bandwidth bottlenecks and improve kernel efficiency.
- Contribute to infrastructure stability and scalability, ensuring reproducibility, consistency across precision formats, and high...
Browse similar: AI jobs · AI startup jobs · Startup jobs · San Francisco Bay Area