# Member of Technical Staff — Model Optimization and Inference (Experienced) at Nuance Labs

- Company: Nuance Labs
- What the company does: We are building visual conversational AI that feels human. Backed by Accel, Lightspeed and South Park Commons.
- Company website: https://www.nuancelabs.ai/
- Type: Startups (AI role)
- Level: Senior
- Location: Seattle, Washington
- Work setup: On-site
- Pay: $250K to $350K base salary per year (USD)
- Posted: 2026-06-05
- Apply by: 2026-10-08
- Apply: https://job-boards.greenhouse.io/nuancelabs/jobs/4277592009
- Page: https://www.1752.vc/careers/jobs/nuance-labs-member-of-technical-staff-model-optimization-and-inference-experienc/

## About the role

We can train a great model. The next problem is making it fast enough to actually use in a real-time conversation — and that gap is enormous. A model that responds in 3 seconds is a demo. A model that responds in under 500ms is a product.

## What they're looking for

- Significant hands-on experience with LLM inference optimization — you’ve shipped work on KV caching, memory layout, attention kernels, or batching strategies in a production or high-traffic research context
- Proven proficiency with inference serving frameworks — vLLM, SGLang, TensorRT-LLM, or similar — including going well beyond default configurations and adapting them to non-standard workloads
- Experience optimizing diffusion model inference (latency reduction, step distillation, caching, or kernel-level work)
- Strong Python and PyTorch skills, comfort reading and writing CUDA or Triton kernels is a significant plus
- A systematic approach to profiling and optimization — you measure first, then optimize
- Familiarity with speculative decoding or other inference-time acceleration techniques

Tags: Research
