Startups · AI

Machine Learning Engineer (Video Understanding & Segmentation)

Maxinsights · Santa Clara · On-site

← All jobs
About Maxinsights

Powering generalist robotics and world models with multi-million-hour annotated egocentric data, hand tracking, upper-body/whole-body motion capture, tactile sensing, and simulation for OpenAI, Google DeepMind, Meta, Figure, 1X, Skild, Genesis AI, and Dyna... Backed by South Park Commons.

About the role

Build and optimize video/image embedding pipelines using CLIP-style and other vision-language embedding models to power large-scale, multi-modal video search and retrieval. Develop LLM-based video understanding systems for semantic indexing, summarization, and question-answering over long-form egocentric and third-person video.

What they're looking for

  • MS or PhD in Computer Science, Electrical Engineering, or a related technical field, or equivalent practical experience
  • 3+ years of hands-on experience in computer vision or multi-modal machine learning, with direct experience in video understanding tasks
  • Strong proficiency in Python and PyTorch, with solid software engineering fundamentals
  • Hands-on experience with CLIP or similar vision-language/video embedding models for retrieval or representation learning
  • Experience building or fine-tuning LLM-based systems for video/image understanding (e.g., captioning, video QA, summarization)
  • Familiarity with agentic system design — tool use, multi-step reasoning, and orchestration frameworks (e.g., LangChain, LlamaIndex, or custom agent loops)
More about this role

We are seeking a highly motivated Machine Learning Engineer to join our core research and development team, focused on video understanding and segmentation. In this role, you will build the systems that let us search, decompose, and describe massive volumes of egocentric and human-robot video at scale — turning raw, unstructured footage into structured, searchable, and richly annotated training data. You will work across video/image embedding models, LLM-based video understanding, and agentic pipelines that orchestrate multiple models into end-to-end workflows. This is a foundational role that directly shapes the data quality and scalability of our entire training data platform.

Build and optimize video/image embedding pipelines using CLIP-style and other vision-language embedding models to power large-scale, multi-modal video search and retrieval.

Develop LLM-based video understanding systems for semantic indexing, summarization, and question-answering over long-form egocentric and third-person video.

Design and implement instruction-level and action-level video chunking/segmentation algorithms that decompose long videos into structured, temporally-aligned clips.

Build automated...

Read the full posting on Maxinsights's site ↗

Engineering

Build your edge while you search

Free tools for founders and investors, plus VC Unfiltered, our take on startups, venture and the people who build them.