Startups · AI

Machine Learning Engineer (Egocentric 3D Human Pose)

Maxinsights · Santa Clara · On-site

← All jobs
About Maxinsights

Powering generalist robotics and world models with multi-million-hour annotated egocentric data, hand tracking, upper-body/whole-body motion capture, tactile sensing, and simulation for OpenAI, Google DeepMind, Meta, Figure, 1X, Skild, Genesis AI, and Dyna... Backed by South Park Commons.

About the role

Build 3D body and hand pose estimation models for egocentric video , covering 2D/3D keypoints, parametric body and hand models (SMPL/SMPL-X, MANO), and full-sequence motion recovery from monocular and stereo first-person cameras. Solve the hard cases specific to the egocentric viewpoint — severe self-occlusion, truncated limbs, extreme perspective foreshortening, hand–object interaction, rapid head motion, and rolling-shutter and motion-blur artifacts.

What they're looking for

  • Bachelor’s, Master’s, or PhD in Computer Science, Machine Learning, Computer Vision, Robotics, or a related technical field, or equivalent practical experience
  • 3+ years of experience building and shipping machine learning systems
  • Proven hands-on experience developing and deploying 3D human pose, hand pose, or human motion tracking models from video
  • Working knowledge of multi-view geometry and camera models: projection, calibration, triangulation, rigid-body transforms, and coordinate-frame management
  • Strong proficiency in Python and at least one major deep learning framework (e.g. PyTorch, TensorFlow)
  • Solid understanding of modern deep learning concepts, training workflows, model evaluation, and real-world, production-oriented ML pipelines
More about this role

We are looking for a Machine Learning Engineer to join our core research and development team, focused on recovering accurate 3D human body and hand motion from egocentric (first-person) video.

Human demonstration data is the fuel for robot learning, and the quality of that data is bounded by how well we can reconstruct what the hands and body actually did. In this role, you will own models and pipelines that turn head-mounted and body-mounted camera streams — often wide-FOV, stereo, motion-blurred, and heavily self-occluded — into metrically accurate, temporally stable 3D pose that is directly usable for robot policy training and human-to-robot retargeting.

You will work across the full stack: capture rig and calibration, ground-truth annotation tooling, model training and evaluation, and production deployment at scale. This role suits engineers who are equally comfortable with multi-view geometry and modern deep learning, and who are motivated by hard, measurable accuracy problems on real-world data.

Build 3D body and hand pose estimation models for egocentric video , covering 2D/3D keypoints, parametric body and hand models (SMPL/SMPL-X, MANO), and full-sequence motion recovery...

Read the full posting on Maxinsights's site ↗

EngineeringML & Data

Build your edge while you search

Free tools for founders and investors, plus VC Unfiltered, our take on startups, venture and the people who build them.