# Research Scientist, Video Foundation Models at Cantina

- Company: Cantina
- What the company does: We help you craft breakthrough products and services. Backed by a16z.
- Company website: https://cantina.co
- Type: Startups (AI role)
- Level: Mid level
- Location: California
- Work setup: On-site
- Pay: $200K to $320K base salary per year (USD)
- Posted: 2026-06-18
- Apply by: 2026-10-08
- Apply: https://jobs.ashbyhq.com/cantina/5411c47f-57d3-4e81-a9f6-71859adc15b6
- Page: https://www.1752.vc/careers/jobs/cantina-research-scientist-video-foundation-models/

## About the role

We are building a core team to develop next-generation native video and omni foundation models for multimodal generation and understanding. Our current focus is large-scale video foundation model development, spanning pre-training, continued training, and post-training for high-quality, controllable, consistent, and efficient generation. Our broader roadmap includes reference- and memory-based generation, multimodal understanding and interaction, and joint audio-video generation.

Tags: Research
