The fastest inference for open models. Backed by Y Combinator.
About the role
Develop and optimize high-performance computing kernels, and work across inference engine internals and serving infrastructure Design, deploy, and operate heterogeneous clusters across vendors
More about this role
Inference performance depends on how efficiently models use the underlying hardware. Techniques across kernels, compilers, runtimes, and serving systems can dramatically improve latency and throughput, but these optimizations are difficult, hardware-specific, and slow to reproduce across new accelerators. And none of it counts until it is running in a customer's production traffic.
At Wafer, we are building AI systems that automatically optimize inference workloads across silicon. The goal is fungible token capacity. Any accelerator optimized toward serving inference most efficiently.
Wafer is well funded and serves trillions of tokens a month for mission critical workloads. We serve the highest performance inference to fast-growing AI startups.
Members of Technical Staff build the systems that make that possible and own the customers running on them. There is no separate solutions team, and no layer between you and the workload.
Develop and optimize high-performance computing kernels, and work across inference engine internals and serving infrastructure
Develop AI agents to do autonomous inference engineering
Design, deploy, and operate heterogeneous clusters across vendors
Own...
Browse similar: AI jobs · AI startup jobs · Startup jobs · San Francisco Bay Area