# Staff / Principal Machine Learning Engineer, Serving - USA at Inworld

- Company: Inworld
- What the company does: Realtime TTS and STT models, LLM serving, and the inference behind both, all through modular APIs. Customers cut voice and AI costs 40% to 95% after moving to Inworld. Backed by Kleiner Perkins, Lightspeed and CRV.
- Company website: https://inworld.ai/
- Type: Startups (AI role)
- Level: Principal and up
- Location: Mountain View, California, USA
- Work setup: Remote
- Pay: $270K to $500K base salary per year (USD)
- Posted: 2026-04-07
- Apply by: 2026-10-08
- Apply: https://jobs.ashbyhq.com/inworld-ai/56fc5614-9273-401a-9cc3-92e8caf27635
- Page: https://www.1752.vc/careers/jobs/inworld-staff-principal-machine-learning-engineer-serving-usa/

## About the role

A year ago, reliably working agentic systems and sub-second multimodal inference at scale barely existed. Nobody has a decade of experience here. So we're not screening for a resume template — we're looking for strong people from varied backgrounds who learn fast, thrive in ambiguity, and can show us what they've built, broken, and understood. You don't need all of this. But you need enough to make a case.

## What they're looking for

- You don't need all of this. But you need enough to make a case
- Inference Optimization. Deep understanding of modern serving frameworks and techniques like vLLM or TRT-LLM
- Model Acceleration . Hands-on experience with quantization, distillation, caching strategies , continuous batching, paged attention, and speculative decoding
- High-Performance Systems. Proficiency in C++, CUDA, Rust, or highly optimized Python. You know how to profile code and squeeze every ounce of performance out of NVIDIA GPUs
- Distributed Systems & Scaling. Experience with Kubernetes, Ray, custom load balancing, multi-GPU/multi-node inference, and reliably handling thousands of concurrent connections
- Public work. Non-trivial systems programming projects, open-source contributions to major inference engines, or deep-dive technical write-ups

Tags: ML Engineering
