At Walden Robotics, we envision a world where general-purpose robots dramatically improve the quality of life for all people—supporting us at home, at work, in factories, on farms, and beyond. Backed by E14 Fund.
About the role
You will build the infrastructure and inference-time algorithms that get policies onto robots and keep them running well. As an MTS you own substantial pieces of the deployment and serving stack, ship them to a live fleet, and work on the models themselves: evaluating them on real hardware, diagnosing failures, and turning findings into model improvements.
What they're looking for
- Software Engineering: Strong fundamentals and experience shipping production services or systems
- Ownership: The ability to own components independently and operate them reliably in production
- Full-Stack Debugging: Comfort debugging across the stack, from model behavior and application logic down to hardware limits
- Data-Driven Reasoning : Comfort reasoning about policy behavior from data and metrics rather than intuition alone
- Operational Mindset: A pragmatic, ownership-minded approach to on-call and operational work
- ML Fundamentals : Hands-on experience training and evaluating deep models, with strong engineering fundamentals in a modern ML framework
More about this role
You will build the infrastructure and inference-time algorithms that get policies onto robots and keep them running well. As an MTS you own substantial pieces of the deployment and serving stack, ship them to a live fleet, and work on the models themselves: evaluating them on real hardware, diagnosing failures, and turning findings into model improvements.
- Deployment Components: Implement and own components of the deployment and on-robot inference pipeline.
- Model Evaluation: Design and run closed-loop evaluations of policies on real robots, and own the picture of what each model can and cannot do.
- Model Improvement : Diagnose policy failure modes from field data and fix them through fine-tuning, data, or model changes, working with the pretraining team.
- Field Evaluation & KPIs: Build telemetry, logging, and KPI-tracking that surface how policies behave in the field, and use that data to attribute failures to the model, the data, or the system.
- Reliability & Performance: Improve reliability and performance of real-time inference and rollout tooling.
- Debugging: Debug issues spanning application code, models, and hardware, often under production urgency.
- Collaboration:...
Browse similar: Startup jobs