Startups · AI

Member of Technical Staff - Kernels & GPU Performance

Gimletlabs · San Francisco, CA · On-site

← All jobs
About Gimletlabs

Backed by Menlo and Sapphire.

About the role

As a Member of Technical Staff, you will build and optimize the low-level execution primitives that turn accelerator performance into production inference performance. Rather than optimizing for one hardware architecture, you will work across accelerators with different execution models, memory hierarchies, capabilities, and software stacks. Your work will shape the latency, throughput, and efficiency Gimlet can achieve across established and emerging hardware architectures.

What they're looking for

  • Strong software engineering fundamentals
  • Experience working on performance-critical systems close to hardware
  • Comfort reasoning about low-level execution behavior, memory hierarchies, and performance tradeoffs
  • Bachelor's degree in a relevant field, or an equivalent combination of education, training, and professional experience
More about this role

Gimlet is building the first multi-silicon neocloud designed for fast, efficient AI inference.

We combine large-scale compute infrastructure with an execution platform that partitions AI workloads and maps each stage to the hardware best suited to run it.

We work with foundation labs, hyperscalers, and AI-native companies, giving our team access to technical problems spanning frontier models, production infrastructure, and emerging hardware.

As a Member of Technical Staff, you will build and optimize the low-level execution primitives that turn accelerator performance into production inference performance.

Rather than optimizing for one hardware architecture, you will work across accelerators with different execution models, memory hierarchies, capabilities, and software stacks. Your work will shape the latency, throughput, and efficiency Gimlet can achieve across established and emerging hardware architectures.

You will work close to the hardware across kernel implementation, memory access, execution behavior, profiling, and performance validation. You will develop optimizations that account for differences between accelerator architectures and partner with compiler, ML systems,...

Read the full posting on Gimletlabs's site ↗

Research and Development

Build your edge while you search

Free tools for founders and investors, plus VC Unfiltered, our take on startups, venture and the people who build them.