# Applied Researcher, Audio at Cartesia

- Company: Cartesia
- What the company does: Integrate real-time text-to-speech with Sonic-3.6, Cartesia. Backed by General Catalyst, Index and Kleiner Perkins.
- Company website: https://www.cartesia.ai/
- Type: Startups (AI role)
- Level: Mid level
- Location: *HQ - San Francisco, CA
- Work setup: On-site
- Pay: $200K to $350K base salary per year (USD)
- Posted: 2025-09-16
- Apply by: 2026-10-08
- Apply: https://jobs.ashbyhq.com/cartesia/d26f0158-f7d1-440b-9c1a-f9b7a884437c
- Page: https://www.1752.vc/careers/jobs/cartesia-applied-researcher-audio/

## About the role

You will be responsible for leading research at the frontier of realtime conversation and human AI interaction. You will contribute across the stack to novel architectures, data, and evals for realtime audio, and translate them into state-of-the-art models used in voice agents around the world. Architect and develop new architectures for realtime audio understanding, generation, and speech-to-speech models that reason jointly over multiple modalities in realtime

## What they're looking for

- Strong applied mindset and ability to balance scientific novelty with product impact
- Excited and able to work across the stack from infra, to data, to evals, to architecture to solve customer problems and build state-of-the-art models
- Deep expertise in deep generative modeling. Previous experience in audio understanding, audio generation, speech-to-speech, or language modeling preferred but not required
- Experience with large-scale training, GPU/TPU acceleration, and model optimization
- Note: Cartesia participates in E-Verify and will provide the federal government with Form I-9 information to confirm employment eligibility after hire

Tags: Research
