Startups · AI

Research Scientist (diffusion)

Genmo · San Francisco HQ · On-site

← All jobs
About Genmo

Genmo is a research lab dedicated to building open, state-of-the-art models for video generation towards unlocking the right brain of AGI. Create high-quality videos with Mochi. Backed by NEA.

About the role

We are seeking an exceptional Research Scientist to join our team, focusing on developing cutting-edge diffusion models for text-to-video generation. In this role, you will be at the forefront of innovation, creating novel architectures and algorithms that transform written descriptions into stunning, coherent video content. Lead research initiatives in advanced diffusion models for text-to-video generation, focusing on improving visual quality, temporal consistency, and semantic fidelity

What they're looking for

  • Ph.D. in Computer Science, Artificial Intelligence, Machine Learning, or a closely related field
  • Strong publication record in top-tier conferences (e.g., CVPR, ICCV, NeurIPS, ICML) with a focus on generative models, particularly diffusion models
  • Extensive experience implementing and optimizing large-scale generative models for image or video tasks
  • Deep understanding of state-of-the-art techniques in text-to-image and text-to-video generation
  • Proficiency in Python and deep learning frameworks such as PyTorch or TensorFlow
  • Excellent communication skills with the ability to explain complex technical concepts to diverse audiences
More about this role

We are Genmo, a research lab developing the world’s most sophisticated video world models to understand, simulate, and interact with the physical world. Our mission is to unlock the right brain of AGI. Join us in advancing physical intelligence and enabling robots to learn and act in a changing world.

We are seeking an exceptional Research Scientist to join our team, focusing on developing cutting-edge diffusion models for text-to-video generation. In this role, you will be at the forefront of innovation, creating novel architectures and algorithms that transform written descriptions into stunning, coherent video content.

Lead research initiatives in advanced diffusion models for text-to-video generation, focusing on improving visual quality, temporal consistency, and semantic fidelity

Develop and implement state-of-the-art algorithms for translating textual descriptions into dynamic video content

Design and conduct rigorous experiments to validate new ideas and evaluate model performance

Collaborate with cross-functional teams to integrate research breakthroughs into our production pipeline

Stay at the cutting edge of the field by regularly reviewing academic literature and...

Read the full posting on Genmo's site ↗

Research

Build your edge while you search

Free tools for founders and investors, plus VC Unfiltered, our take on startups, venture and the people who build them.