Backed by SOSV.
About the role
Most teams building agents make design decisions by intuition and anecdote. Someone tries a new memory scheme, it feels better, it ships. We think that is the central failure of the field right now, and we are building this team to work the other way: every decision about how our agent systems are constructed should be settled by evidence.
What they're looking for
- Graduate degree in CS, ML, statistics, or a related field, or equivalent research experience. We care about demonstrated research judgment, not credentials
- Fluent in the current literature and able to judge it. You read papers continuously, can tell a real result from a well-marketed one, and have opinions about which recent directions are overrated
- Deep understanding of how LLMs work — pretraining through the post-training stack, and what actually happens at inference. You reason from mechanism, not just from published numbers
- Firm grasp of the full evaluation pipeline : sourcing eval data, constructing the loop, automating the climb. Having done all three for a real system, rather than one in isolation, is the strongest signal for this role
- Rigorous experimentalist. You design experiments that can fail, you understand variance and power, and you are comfortable saying an intervention didn't work
- A real engineering background. Strong Python, comfortable with production systems and data, able to stand up the infrastructure your own experiment needs
More about this role
Nuro is a self-driving technology company on a mission to make autonomy accessible to all. Founded in 2016, Nuro is building the world’s most scalable driver, combining cutting-edge AI with automotive-grade hardware. Nuro licenses its core technology, the Nuro Driver™, to support a wide range of applications, from robotaxis and commercial fleets to personally owned vehicles. With technology proven over years of self-driving deployments, Nuro gives automakers and mobility platforms a clear path to AVs at commercial scale, empowering a safer, richer, and more connected future.
Frontier models are fungible. Any team can rent the same intelligence we can, and the model we build on today will be replaced within a month. What is not fungible is the infrastructure that decides whether an autonomous system's output can be trusted — evaluation, verification, and the discipline to gate on evidence instead of impressions. Nuro has spent a decade building exactly that discipline for a robot that drives on public roads, and this team turns it inward: we build the platform that lets AI agents operate autonomously inside Nuro's own engineering organization, under the same standard of proof we...
Browse similar: AI jobs · AI startup jobs · Startup jobs · San Francisco Bay Area