# ML Model Serving Engineer at Sesame AI

- Company: Sesame AI
- What the company does: Sesame builds personal agents for curious people. Follow a thought. Work out an idea. Discover something new. Preview available now on iOS. Coming to intelligent eyewear in 2027. Backed by Sequoia and Redpoint.
- Company website: https://www.sesame.com/
- Type: Startups (AI role)
- Level: Mid level
- Location: San Francisco
- Work setup: On-site
- Pay: $175K to $280K base salary per year (USD)
- Posted: 2025-03-14
- Apply by: 2026-10-08
- Apply: https://jobs.ashbyhq.com/sesame/35793528-2b5c-47b3-9422-eaced2b69f63
- Page: https://www.1752.vc/careers/jobs/sesame-ai-ml-model-serving-engineer/

## About the role

Turbocharge our serving layer, consisting of a variety of LLM, speech, and vision models. Partner with ML infrastructure and training engineers to build a fast, cost-effective, accurate, and reliable serving layer to power a new consumer product category.

## What they're looking for

- Expert in some differentiable array computing framework, preferably PyTorch
- Expert in optimizing machine learning models for serving reliably at high throughput, with low latency
- Significant systems programming experience, ex. Experience working on high-performance server systems—you’d be just as comfortable with the internals of VLLM as you would with a complex PyTorch codebase
- Significant performance engineering experience, ex. Bottleneck analysis in high-scale server systems or profiling low-level systems code
- Always up to date on the latest techniques for model serving optimization
- Familiarity with high-performance LLM serving, ex. experience with VLLM, SGlang deployment, and internals

Tags: Software
