# Research Engineer, Data Infrastructure (Language Modeling) at Cartesia

- Company: Cartesia
- What the company does: Integrate real-time text-to-speech with Sonic-3.6, Cartesia. Backed by General Catalyst, Index and Kleiner Perkins.
- Company website: https://www.cartesia.ai/
- Type: Startups (AI role)
- Level: Mid level
- Location: *HQ - San Francisco, CA
- Work setup: On-site
- Pay: $200K to $350K base salary per year (USD)
- Posted: 2026-09-24
- Apply by: 2026-11-08
- Apply: https://jobs.ashbyhq.com/cartesia/6ca9b352-6a7b-42a3-a7c4-8f071712db90
- Page: https://www.1752.vc/careers/jobs/cartesia-research-engineer-data-infrastructure-language-modeling/

## About the role

Data is the lifeblood of our models, and we are looking for a Research Engineer, Data Infrastructure to build the datasets and systems that power pretraining at Cartesia. In this role, you will write performant, scalable infrastructure to acquire, process, and curate massive datasets, and partner closely with research to optimize the characteristics and composition of data mixtures. Your work will directly shape the capabilities and quality of our foundational models.

## What they're looking for

- Hands-on experience with ML data infrastructure: training data pipelines, dataset versioning, large-scale data loading, and the interplay between data systems and model training and inference
- Strong modern engineering execution: clean, well-tested code, fluency with current tools, and a willingness to pick the right tool for the problem rather than defaulting to familiar patterns
- Familiarity with building and evaluating datasets for generative models and reasonable working knowledge of how they're trained and inference

Tags: Research
