# ML Research Engineer (Distributed Training) at Morphic Robot

- Company: Morphic Robot
- What the company does: Morphic AI is transforming the future of storytelling using breakthrough AI technologies. Go from idea to final video effortlessly with Morphic.
- Company website: https://morphic.com
- Type: Startups (AI role)
- Level: Mid level
- Location: Palo Alto
- Work setup: On-site
- Pay: $200K to $280K base salary per year (USD)
- Posted: 2026-03-26
- Apply by: 2026-10-08
- Apply: https://jobs.ashbyhq.com/metamorphic/20d7f6a3-d768-40d5-9d40-84fee852e866
- Page: https://www.1752.vc/careers/jobs/morphic-robot-ml-research-engineer-distributed-training/

## About the role

We are hiring Research Engineers to join our growing AI research team. You will work on building and scaling the distributed systems that enable training Metamorphic’s state-of-the-art foundation models across thousands of GPU’s. This is a high-impact, technically deep role working at the frontier of ML research and engineering.

## What they're looking for

- Bachelor's degree or equivalent experience in Computer Science, Machine Learning, or a related field
- Strong software engineering skills with a proven track record of building complex systems
- Hands-on experience building and debugging distributed training infrastructure (PyTorch FSDP, DeepSpeed ZeRO, Megatron, TorchTitan, or similar) and optimizing advanced parallelism strategies
- Strong understanding of GPU architecture and performance: memory hierarchy, tensor core utilization, bandwidth vs compute limitations
- Strong understanding of the NVIDIA ecosystem: CUDA, NCCL, NVLink/NVSwitch topologies, mixed-precision training (MXFP8/NVFP4), and profiling tools
- Deep familiarity with PyTorch internals, including torch.distributed, autograd, memory management, and torch.compile

Tags: Technical Staff
