# Member of Technical Staff, Multimodal Speech at Hark

- Company: Hark
- What the company does: At Hark, we are building the most advanced personal intelligence in the world.
- Company website: https://hark.com
- Type: Startups
- Level: Senior
- Location: San Jose
- Work setup: On-site
- Pay: $180K to $450K base salary per year (USD)
- Posted: 2026-03-23
- Apply by: 2026-10-08
- Apply: https://job-boards.greenhouse.io/hark/jobs/4194153009
- Page: https://www.1752.vc/careers/jobs/hark-member-of-technical-staff-multimodal-speech/

## About the role

The Omni team at Hark is building the next generation of AI experiences beyond text, enabling models to understand and generate content across multiple modalities, including text, audio. Our goal is to create seamless, real-time multimodal intelligence that powers intuitive and immersive user experiences.

## What they're looking for

- Proven track record of advancing speech or audio models through innovations in data, modeling, or training
- Strong experience in speech/audio domains such as ASR, TTS, speech-to-speech, or audio foundation models
- Experience with large-scale machine learning systems and distributed training
- Strong background in data-driven experimentation, systematic evaluation, and model iteration
- Strong ownership mindset and ability to drive end-to-end impact from research to production

Tags: AI Foundation Models
