Startups · AI

Member of Technical Staff

Wafer · San Francisco · On-site

← All jobs
About Wafer

The fastest inference for open models. Backed by Y Combinator.

About the role

Develop and optimize high-performance computing kernels, and work across inference engine internals and serving infrastructure Design, deploy, and operate heterogeneous clusters across vendors

More about this role

Inference performance depends on how efficiently models use the underlying hardware. Techniques across kernels, compilers, runtimes, and serving systems can dramatically improve latency and throughput, but these optimizations are difficult, hardware-specific, and slow to reproduce across new accelerators. And none of it counts until it is running in a customer's production traffic.

At Wafer, we are building AI systems that automatically optimize inference workloads across silicon. The goal is fungible token capacity. Any accelerator optimized toward serving inference most efficiently.

Wafer is well funded and serves trillions of tokens a month for mission critical workloads. We serve the highest performance inference to fast-growing AI startups.

Members of Technical Staff build the systems that make that possible and own the customers running on them. There is no separate solutions team, and no layer between you and the workload.

Develop and optimize high-performance computing kernels, and work across inference engine internals and serving infrastructure

Develop AI agents to do autonomous inference engineering

Design, deploy, and operate heterogeneous clusters across vendors

Own...

Read the full posting on Wafer's site ↗

Engineering

Build your edge while you search

Free tools for founders and investors, plus VC Unfiltered, our take on startups, venture and the people who build them.