# Machine Learning Engineer, Speech - Joint Audio-Video Modeling at Cantina

- Company: Cantina
- What the company does: We help you craft breakthrough products and services. Backed by a16z.
- Company website: https://cantina.co
- Type: Startups (AI role)
- Level: Mid level
- Location: Remote (U.S. or Europe)
- Work setup: Remote
- Pay: $200K to $220K base salary per year (USD)
- Posted: 2026-07-30
- Apply by: 2026-10-08
- Apply: https://jobs.ashbyhq.com/cantina/0b1ec7c7-ca1f-4d86-9242-f6397f60ba34
- Page: https://www.1752.vc/careers/jobs/cantina-machine-learning-engineer-speech-joint-audio-video-modeling/

## About the role

We're looking for a Research / ML Engineer to join our Speech Team to build state-of-the-art speech and audio generation systems end-to-end from data specs through production inference with a focus on joint audio-video modeling.

Tags: Research
