Startups · AI

Member of Technical Staff - Multilingual Data

Reflection AI · San Francisco, CA · On-site

← All jobs
About Reflection AI

Make intelligence open and accessible to all. Backed by Battery, Lightspeed and Sequoia.

About the role

Design and operate large-scale multilingual data pipelines — sourcing, cleaning, deduplication, language identification, and script normalization across high- and low resource languages. Define and enforce quality bars for multilingual corpora, including translation quality, cultural fidelity, toxicity, and contamination checks.

More about this role

Reflection is a research lab making intelligence open and accessible for everyone to use, customize, and build on. We build open models that let anyone control their intelligence and help shape the future of AI. Our mission: make intelligence open and accessible to all.

Design and operate large-scale multilingual data pipelines — sourcing, cleaning, deduplication, language identification, and script normalization across high- and low resource languages.

Define and enforce quality bars for multilingual corpora, including translation quality, cultural fidelity, toxicity, and contamination checks.

Design and run scientific experiments to advance our understanding of scaling large language models to improve multilingual data efficiency.

Lead small research projects independently while collaborating on larger initiatives.

Build evaluation sets and diagnostics that expose where model behavior degrades by language, register, or domain, and close those gaps with targeted data.

Work with pre-training, mid-training, and post-training teams to land measurable, step-function improvements in multilingual capability.

Strong software engineering fundamentals and comfort processing web-scale...

Read the full posting on Reflection AI's site ↗

Research

Build your edge while you search

Free tools for founders and investors, plus VC Unfiltered, our take on startups, venture and the people who build them.