Startups · AI

Research Member of Technical Staff- Post-training & Robot Learning

Rhoda · Mountain View · On-site

← All jobs
About Rhoda

Redefining Robotic Intelligence. Backed by Khosla.

About the role

Design and implement RL training pipelines to improve robot policy performance beyond what imitation learning alone achieves — reward design, online data collection, and policy optimization Develop and apply RL algorithms (PPO, GRPO, or similar) adapted to the video prediction setting, including reward modeling and feedback collection strategies for physical task performance

What they're looking for

  • Hands-on experience with robot systems, robotic policy learning, or autonomous systems in an industry or research setting (robotics, self-driving, or similar physical AI domains)
  • Strong understanding of robot policy learning: imitation learning, behavior cloning, and how RL builds on top of it
  • Practical familiarity with real robot hardware, deployment constraints, and sensor modalities (vision, proprioception)
  • Solid ML skills with hands-on PyTorch experience
  • Ability to diagnose policy failures, reason about distribution shift, and iterate effectively on data and training strategies
  • Comfort with ambiguity and fast-changing research priorities
More about this role

At Rhoda AI, we’re building the next generation of generalist intelligent robots. We own the full robotics stack from high-performance hardware and robot systems to the infrastructure and state-of-the-art foundation world models that control our robots. Our robots are designed to be generalists capable of operating in complex, real-world environments and handling long-tail edge cases, made possible by our cutting edge research and end-to-end system design. We've raised over $450M and are investing aggressively in model research, infrastructure, hardware development, and manufacturing scale-up to make generalist robotics a reality.

We're looking for Research Scientists and Research Engineers with deep robotics or autonomous systems domain knowledge to adapt our web-pretrained video model to real robot tasks. Post-training at Rhoda means taking a causal video generation model pretrained on internet-scale data and fine-tuning it on robot-collected demonstrations to produce reliable, generalizable behavior — with as little task-specific data as possible. We hire across levels — from senior to staff.

Design and implement RL training pipelines to improve robot policy performance beyond...

Read the full posting on Rhoda's site ↗

Research

Build your edge while you search

Free tools for founders and investors, plus VC Unfiltered, our take on startups, venture and the people who build them.