Make intelligence open and accessible to all. Backed by Battery, Lightspeed and Sequoia.
About the role
Design and operate large-scale multilingual data pipelines — sourcing, cleaning, deduplication, language identification, and script normalization across high- and low resource languages. Define and enforce quality bars for multilingual corpora, including translation quality, cultural fidelity, toxicity, and contamination checks.
More about this role
Reflection is a research lab making intelligence open and accessible for everyone to use, customize, and build on. We build open models that let anyone control their intelligence and help shape the future of AI. Our mission: make intelligence open and accessible to all.
Design and operate large-scale multilingual data pipelines — sourcing, cleaning, deduplication, language identification, and script normalization across high- and low resource languages.
Define and enforce quality bars for multilingual corpora, including translation quality, cultural fidelity, toxicity, and contamination checks.
Design and run scientific experiments to advance our understanding of scaling large language models to improve multilingual data efficiency.
Lead small research projects independently while collaborating on larger initiatives.
Build evaluation sets and diagnostics that expose where model behavior degrades by language, register, or domain, and close those gaps with targeted data.
Work with pre-training, mid-training, and post-training teams to land measurable, step-function improvements in multilingual capability.
Strong software engineering fundamentals and comfort processing web-scale...
Browse similar: AI jobs · AI startup jobs · Startup jobs · San Francisco Bay Area