Startups · AI

Member of Technical Staff, Vision / Language

XDOF · San Mateo Hybrid · Remote

← All jobs
About XDOF

Defining motion for autonomous systems. Backed by a16z.

About the role

Design and implement vision-language pipelines for egocentric and teleoperation video: structured captioning, temporal grounding, action-conditioned scene understanding, and semantic annotation at scale Develop and evaluate representations that bridge visual perception, language, and low-level robot action — spanning VLAs, video prediction, and world models

What they're looking for

  • MS or PhD in Computer Science, Robotics, Machine Learning, or a related field from a top-tier program
  • 3–7 years of research or applied research experience (industry or academic) in one or more of: vision-language models, video understanding, robot learning, or generative modeling
  • Deep fluency in PyTorch, working knowledge of large-scale training infrastructure (distributed training, mixed precision, large batch workflows)
  • Published work or demonstrable impact in VLMs/VLAs, video representation learning, imitation learning, or a closely related area
  • Strong engineering fundamentals — you can design clean systems, not just run experiments
More about this role

Frontier labs are racing to build general-purpose robots, and the bottleneck isn't compute. It's data. At XDOF, we're building the foundation behind the foundation models: the data collection systems, annotation pipelines, exabyte-scale data infrastructure, and software toolchain that enable our partners to push the field forward.

We're hiring a Research Engineer / Scientist to help lead technical efforts at the intersection of vision-language models and robot learning. You will build systems that turn raw egocentric and teleoperation video into high-signal training data for VLA models, and increasingly, contribute to the models themselves.

Beyond pipelines, you will drive research into what makes robot data useful : discovering new metadata (contact events, affordance labels, implicit reward signals, dynamics priors from video) that unlock capabilities current approaches miss. You'll explore how structured annotations can improve cross-embodiment transfer, automatic curriculum generation, and world models that predict what actually matters for manipulation. The data layer isn't downstream of the research. It is the research.

Design and implement vision-language pipelines for...

Read the full posting on XDOF's site ↗

Robotics

Build your edge while you search

Free tools for founders and investors, plus VC Unfiltered, our take on startups, venture and the people who build them.