Backed by Menlo and Sapphire.
About the role
As a Member of Technical Staff focused on ML Systems, you will build the inference systems that execute models end-to-end in production. You will work on the systems that determine how inference executes across that pipeline: how requests are batched and scheduled, how stages are placed and scaled, how KV cache and intermediate state move between accelerators, and how the system balances latency, throughput, and utilization across different hardware characteristics.
What they're looking for
- Strong software engineering fundamentals
- Experience building or operating ML inference or model serving systems
- Comfort reasoning about performance, memory usage, and system behavior under load
- Bachelor's degree in a relevant field, or an equivalent combination of education, training, and professional experience
More about this role
Gimlet is building the first multi-silicon neocloud designed for fast, efficient AI inference.
We combine large-scale compute infrastructure with an execution platform that partitions AI workloads and maps each stage to the hardware best suited to run it.
We work with foundation labs, hyperscalers, and AI-native companies, giving our team access to technical problems spanning frontier models, production infrastructure, and emerging hardware.
As a Member of Technical Staff focused on ML Systems, you will build the inference systems that execute models end-to-end in production.
You will work on the systems that determine how inference executes across that pipeline: how requests are batched and scheduled, how stages are placed and scaled, how KV cache and intermediate state move between accelerators, and how the system balances latency, throughput, and utilization across different hardware characteristics.
You will work across model serving, batching, scheduling, concurrency, KV cache management, and memory placement. You will help bring up models on novel hardware. You will support new model architectures and inference techniques, improve performance under real production workloads,...
Browse similar: AI jobs · AI startup jobs · Startup jobs · San Francisco Bay Area