# Member of Technical Staff, Vision / Language at XDOF

- Company: XDOF
- What the company does: Defining motion for autonomous systems. Backed by a16z.
- Company website: https://www.xdof.ai
- Type: Startups (AI role)
- Level: Senior
- Location: San Mateo Hybrid
- Work setup: Remote
- Posted: 2026-06-05
- Apply by: 2026-10-08
- Apply: https://jobs.ashbyhq.com/xdof/1e7de4e3-094f-4267-876d-2ccb8b2246eb
- Page: https://www.1752.vc/careers/jobs/xdof-member-of-technical-staff-vision-language/

## About the role

Design and implement vision-language pipelines for egocentric and teleoperation video: structured captioning, temporal grounding, action-conditioned scene understanding, and semantic annotation at scale Develop and evaluate representations that bridge visual perception, language, and low-level robot action — spanning VLAs, video prediction, and world models

## What they're looking for

- MS or PhD in Computer Science, Robotics, Machine Learning, or a related field from a top-tier program
- 3–7 years of research or applied research experience (industry or academic) in one or more of: vision-language models, video understanding, robot learning, or generative modeling
- Deep fluency in PyTorch, working knowledge of large-scale training infrastructure (distributed training, mixed precision, large batch workflows)
- Published work or demonstrable impact in VLMs/VLAs, video representation learning, imitation learning, or a closely related area
- Strong engineering fundamentals — you can design clean systems, not just run experiments

Tags: Robotics
