# Senior Inference Optimization ML Engineer at Rhoda

- Company: Rhoda
- What the company does: Redefining Robotic Intelligence. Backed by Khosla.
- Company website: https://www.rhoda.ai/
- Type: Startups (AI role)
- Level: Senior
- Location: Mountain View
- Work setup: On-site
- Pay: $175K to $250K base salary per year (USD)
- Posted: 2026-05-12
- Apply by: 2026-10-08
- Apply: https://jobs.ashbyhq.com/rhoda-ai/5fbe9c15-342b-4d46-b1bf-34d99d2f243c
- Page: https://www.1752.vc/careers/jobs/rhoda-senior-inference-optimization-ml-engineer/

## About the role

Own inference performance end-to-end — diagnose and improve latency, throughput, and efficiency of large foundation models in production Build systematic performance attribution: latency decomposition (compute vs. memory bandwidth vs. I/O), bottleneck identification, and prioritization across model families

## What they're looking for

- 3+ years of experience in inference optimization, ML systems, or a closely related field
- Deep hands-on experience with modern ML stacks (PyTorch required, JAX a plus)
- Strong understanding of compute, memory bandwidth, and I/O bottlenecks in large model inference
- Experience with model optimization techniques: quantization (INT8/FP8/AWQ), distillation, pruning, and compilation
- Familiarity with inference serving frameworks (e.g., Triton, TensorRT, vLLM, TorchServe)
- Exceptional debugging and measurement ability: turn "inference is slow" into clear bottlenecks, experiments, and validated improvements

Tags: Software
