Genmo is a research lab dedicated to building open, state-of-the-art models for video generation towards unlocking the right brain of AGI. Create high-quality videos with Mochi. Backed by NEA.
About the role
We are seeking an exceptional Research Scientist to join our team, focusing on developing cutting-edge diffusion models for text-to-video generation. In this role, you will be at the forefront of innovation, creating novel architectures and algorithms that transform written descriptions into stunning, coherent video content. Lead research initiatives in advanced diffusion models for text-to-video generation, focusing on improving visual quality, temporal consistency, and semantic fidelity
What they're looking for
- Ph.D. in Computer Science, Artificial Intelligence, Machine Learning, or a closely related field
- Strong publication record in top-tier conferences (e.g., CVPR, ICCV, NeurIPS, ICML) with a focus on generative models, particularly diffusion models
- Extensive experience implementing and optimizing large-scale generative models for image or video tasks
- Deep understanding of state-of-the-art techniques in text-to-image and text-to-video generation
- Proficiency in Python and deep learning frameworks such as PyTorch or TensorFlow
- Excellent communication skills with the ability to explain complex technical concepts to diverse audiences
More about this role
We are Genmo, a research lab developing the world’s most sophisticated video world models to understand, simulate, and interact with the physical world. Our mission is to unlock the right brain of AGI. Join us in advancing physical intelligence and enabling robots to learn and act in a changing world.
We are seeking an exceptional Research Scientist to join our team, focusing on developing cutting-edge diffusion models for text-to-video generation. In this role, you will be at the forefront of innovation, creating novel architectures and algorithms that transform written descriptions into stunning, coherent video content.
Lead research initiatives in advanced diffusion models for text-to-video generation, focusing on improving visual quality, temporal consistency, and semantic fidelity
Develop and implement state-of-the-art algorithms for translating textual descriptions into dynamic video content
Design and conduct rigorous experiments to validate new ideas and evaluate model performance
Collaborate with cross-functional teams to integrate research breakthroughs into our production pipeline
Stay at the cutting edge of the field by regularly reviewing academic literature and...
Browse similar: AI jobs · AI startup jobs · Startup jobs · San Francisco Bay Area