Startups · AI

Senior Inference Optimization ML Engineer

Rhoda · Mountain View · On-site

← All jobs
About Rhoda

Redefining Robotic Intelligence. Backed by Khosla.

About the role

Own inference performance end-to-end — diagnose and improve latency, throughput, and efficiency of large foundation models in production Build systematic performance attribution: latency decomposition (compute vs. memory bandwidth vs. I/O), bottleneck identification, and prioritization across model families

What they're looking for

  • 3+ years of experience in inference optimization, ML systems, or a closely related field
  • Deep hands-on experience with modern ML stacks (PyTorch required, JAX a plus)
  • Strong understanding of compute, memory bandwidth, and I/O bottlenecks in large model inference
  • Experience with model optimization techniques: quantization (INT8/FP8/AWQ), distillation, pruning, and compilation
  • Familiarity with inference serving frameworks (e.g., Triton, TensorRT, vLLM, TorchServe)
  • Exceptional debugging and measurement ability: turn "inference is slow" into clear bottlenecks, experiments, and validated improvements
More about this role

At Rhoda AI, we’re building the next generation of generalist intelligent robots. We own the full robotics stack from high-performance hardware and robot systems to the infrastructure and state-of-the-art foundation world models that control our robots. Our robots are designed to be generalists capable of operating in complex, real-world environments and handling long-tail edge cases, made possible by our cutting edge research and end-to-end system design. We've raised over $450M and are investing aggressively in model research, infrastructure, hardware development, and manufacturing scale-up to make generalist robotics a reality.

We're looking for an Inference Optimization MLE to help build and operate the systems that make our foundation models run fast and efficiently in production. You'll be responsible for squeezing maximum performance out of large multimodal models, across cloud and on-robot deployment targets. You will working closely with research and robotics teams to close the gap between training and real-world deployment.

Own inference performance end-to-end — diagnose and improve latency, throughput, and efficiency of large foundation models in production

Build...

Read the full posting on Rhoda's site ↗

Software

Build your edge while you search

Free tools for founders and investors, plus VC Unfiltered, our take on startups, venture and the people who build them.