# Member of Technical Staff — Model Optimization and Inference (New Grad) at Nuance Labs

- Company: Nuance Labs
- What the company does: We are building visual conversational AI that feels human. Backed by Accel, Lightspeed and South Park Commons.
- Company website: https://www.nuancelabs.ai/
- Type: Startups (AI role)
- Level: Senior
- Location: Seattle, Washington
- Work setup: On-site
- Pay: $200K to $300K base salary per year (USD)
- Posted: 2026-06-11
- Apply by: 2026-10-08
- Apply: https://job-boards.greenhouse.io/nuancelabs/jobs/4283935009
- Page: https://www.1752.vc/careers/jobs/nuance-labs-member-of-technical-staff-model-optimization-and-inference-new-grad/

## About the role

We can train a great model. The next problem is making it fast enough to actually use in a real-time conversation — and that gap is enormous. A model that responds in 3 seconds is a demo. A model that responds in under 500ms is a product.

## What they're looking for

- BS, MS, or PhD in CS, ML, or a related field — completed or in the final stretch
- Exposure to inference serving frameworks (vLLM, SGLang, TensorRT-LLM, or similar) — even at a research or hobby level
- Strong Python and PyTorch skills, familiarity with CUDA or Triton is a significant plus
- A systematic approach to profiling and optimization — you measure first, then optimize
- Curiosity about diffusion inference, speculative decoding, quantization, or other inference-time acceleration techniques

Tags: Research
