Startups · AI

Research Member of Technical Staff- Efficient Modeling

Rhoda · Mountain View · On-site

← All jobs
About Rhoda

Redefining Robotic Intelligence. Backed by Khosla.

About the role

Research and implement model compression techniques: quantization, pruning, structured sparsity, distillation, and low-rank approximation Design efficient architectures and attention mechanisms suited to real-time inference on edge and robot hardware

What they're looking for

  • Strong understanding of model compression and efficient architectures for large models
  • Hands-on experience with quantization, distillation, or pruning applied to transformers or large neural networks
  • Deep knowledge of where efficiency gains are possible in modern architectures
  • Proficiency with PyTorch and familiarity with hardware-aware optimization (CUDA, TensorRT, or similar)
  • Ability to run principled experiments that characterize capability-efficiency tradeoffs
  • PhD in ML, CS, or a related field — or equivalent research/engineering experience
More about this role

At Rhoda AI, we’re building the next generation of generalist intelligent robots. We own the full robotics stack from high-performance hardware and robot systems to the infrastructure and state-of-the-art foundation world models that control our robots. Our robots are designed to be generalists capable of operating in complex, real-world environments and handling long-tail edge cases, made possible by our cutting edge research and end-to-end system design. We've raised over $450M and are investing aggressively in model research, infrastructure, hardware development, and manufacturing scale-up to make generalist robotics a reality.

We're looking for a Research Scientist or Research Engineer focused on model efficiency — making our foundation world models faster, smaller, and more deployable without sacrificing capability. This work is critical to closing the gap between research-scale models and real-time operation on robot hardware.

Research and implement model compression techniques: quantization, pruning, structured sparsity, distillation, and low-rank approximation

Design efficient architectures and attention mechanisms suited to real-time inference on edge and robot...

Read the full posting on Rhoda's site ↗

Research

Build your edge while you search

Free tools for founders and investors, plus VC Unfiltered, our take on startups, venture and the people who build them.