# Machine Learning Engineer (Video Understanding & Segmentation) at Maxinsights

- Company: Maxinsights
- What the company does: Powering generalist robotics and world models with multi-million-hour annotated egocentric data, hand tracking, upper-body/whole-body motion capture, tactile sensing, and simulation for OpenAI, Google DeepMind, Meta, Figure, 1X, Skild, Genesis AI, and Dyna... Backed by South Park Commons.
- Company website: https://maxinsights.ai
- Type: Startups (AI role)
- Level: Mid level
- Location: Santa Clara
- Work setup: On-site
- Posted: 2026-09-25
- Apply by: 2026-11-09
- Apply: https://jobs.ashbyhq.com/maxinsights/631ee3d8-07ca-455d-b6a2-dee4d3d33a72
- Page: https://www.1752.vc/careers/jobs/maxinsights-machine-learning-engineer-video-understanding-and-segmentation/

## About the role

Build and optimize video/image embedding pipelines using CLIP-style and other vision-language embedding models to power large-scale, multi-modal video search and retrieval. Develop LLM-based video understanding systems for semantic indexing, summarization, and question-answering over long-form egocentric and third-person video.

## What they're looking for

- MS or PhD in Computer Science, Electrical Engineering, or a related technical field, or equivalent practical experience
- 3+ years of hands-on experience in computer vision or multi-modal machine learning, with direct experience in video understanding tasks
- Strong proficiency in Python and PyTorch, with solid software engineering fundamentals
- Hands-on experience with CLIP or similar vision-language/video embedding models for retrieval or representation learning
- Experience building or fine-tuning LLM-based systems for video/image understanding (e.g., captioning, video QA, summarization)
- Familiarity with agentic system design — tool use, multi-step reasoning, and orchestration frameworks (e.g., LangChain, LlamaIndex, or custom agent loops)

Tags: Engineering
