Snorkel AI builds specialized training data, benchmarks, and evaluation environments that help frontier models and agents perform in high-stakes domains. Backed by Greylock, Lightspeed and GV.
About the role
We're looking for a Research Scientist to lead the design of next-generation benchmarks and datasets that push the boundaries of frontier model evaluation. You'll define what "good" looks like across a range of hard tasks, drawing on conversations with customers and academic partners to ground your datasets in real performance gaps.
What they're looking for
- Strong research background in AI/ML evaluation, NLP, or related fields, with a track record of rigorous experimental design — especially around measuring the impact of training and evaluation data on model behavior
- Exceptional communication skills — able to present complex technical findings clearly to both technical and non-technical audiences
- Comfort operating in a fast-moving, cross-functional environment with ambiguous problem spaces
- Genuine interest in GTM strategy, startup dynamics, and the commercial side of AI data services
- Ph.D. in machine learning, NLP, or a related field preferred, equivalent industry or research lab experience considered
More about this role
At Snorkel, we believe meaningful AI doesn’t start with the model, it starts with the data.
We’re on a mission to help enterprises transform expert knowledge into specialized AI at scale. The AI landscape has gone through incredible changes since 2015, when Snorkel started as a research project in the Stanford AI Lab, to the generative AI breakthroughs of today. But one thing has remained constant: the data you use to build AI is the key to achieving differentiation, high performance, and production-ready systems. We work with some of the world’s largest organizations to empower scientists, engineers, financial experts, product creators, journalists, and more to build custom AI with their data faster than ever before. Excited to help us redefine how AI is built? Apply to be the newest Snorkeler!
We're looking for a Research Scientist to lead the design of next-generation benchmarks and datasets that push the boundaries of frontier model evaluation. You'll define what "good" looks like across a range of hard tasks, drawing on conversations with customers and academic partners to ground your datasets in real performance gaps. You'll build both the benchmarks that measure those gaps...
Browse similar: AI jobs · AI startup jobs · Startup jobs · Remote jobs · San Francisco Bay Area