Startups

GPU Kernel Engineer

Sciforium · San Francisco, CA · Remote

← All jobs
About Sciforium

Sciforium builds the next generation of AI models with unprecedented efficiency, privacy, and versatility. Backed by SignalFire.

About the role

We are seeking a highly skilled GPU Kernel Engineer who is passionate about pushing the limits of performance on modern accelerators. In this role, you will design and optimize custom GPU kernels that power next-generation large-scale AI systems. You will work across the hardware–software stack, from low-level kernel development to integrating optimized ops into high-level ML frameworks used for large-scale training and inference.

What they're looking for

  • 5+ years of industry or research experience in GPU kernel development or high-performance computing
  • Bachelor’s, Master’s, or PhD in Computer Science, Computer Engineering, Electrical Engineering, Applied Mathematics, or a related field
  • Strong programming skills in C++, Python, and familiarity with ML frameworks
  • Deep expertise in CUDA/ROCm, GPU memory models, and performance optimization strategies
  • Hands-on experience with Triton and/or JAX Pallas for custom kernel development
  • Strong understanding of PTX, GPU ASM, and low-level GPU execution
More about this role

Sciforium is an AI infrastructure company developing next-generation multimodal AI models and a proprietary, high-efficiency serving platform. Backed by multi-million-dollar funding and direct sponsorship from AMD with hands-on support from AMD engineers the team is scaling rapidly to build the full stack powering frontier AI models and real-time applications.

We are seeking a highly skilled GPU Kernel Engineer who is passionate about pushing the limits of performance on modern accelerators. In this role, you will design and optimize custom GPU kernels that power next-generation large-scale AI systems. You will work across the hardware–software stack, from low-level kernel development to integrating optimized ops into high-level ML frameworks used for large-scale training and inference.

This role is ideal for someone who thrives at the intersection of GPU programming, systems engineering, and cutting-edge AI workloads, and who wants to make meaningful contributions to the efficiency and scalability of our ML platform.

Design, implement, and optimize custom GPU kernels using C++, PTX, CUDA, ROCm, Triton, and/or JAX Pallas.

Profile and optimize end-to-end performance of ML...

Read the full posting on Sciforium's site ↗

Engineering

Build your edge while you search

Free tools for founders and investors, plus VC Unfiltered, our take on startups, venture and the people who build them.