Connectionism: Research Blog by Thinking Machines Lab. Backed by a16z, Accel and GV.
About the role
The role of pre-training researchers sits at the core of our roadmap. This work blends research with large-scale data engineering to help assemble the pre-training datasets and data systems that underpin the next generation of AI models. You’ll design and implement methods for sourcing, curating, and analyzing pre-training data for quality and performance.
What they're looking for
- Proficiency in Python and familiarity with at least one deep learning framework (e.g., PyTorch, TensorFlow, or JAX). Comfortable with debugging distributed training and writing code that scales
- Bachelor’s degree or equivalent experience in Computer Science, Machine Learning, Physics, Mathematics, or a related discipline with strong theoretical and empirical grounding
- Clarity in communication, an ability to explain complex technical concepts in writing
- Preferred qualifications — we encourage you to apply even if you don’t meet all preferred qualifications, but at least some:
- A strong grasp of probability, statistics, and ML fundamentals. You can look at experimental data and distinguish between real effects, noise, and bugs
- Experience with curation, preprocessing, and analysis of large-scale text, code, or multimodal datasets
More about this role
The mission of Thinking Machines is to build AI that extends human will and judgment. We are training frontier models with Inkling, developing Tinker to let people make models their own, and crafting interfaces that broaden human-AI communication. We believe the future worth building is human, and we're hiring people who want to build it.
The role of pre-training researchers sits at the core of our roadmap. This work blends research with large-scale data engineering to help assemble the pre-training datasets and data systems that underpin the next generation of AI models. You’ll design and implement methods for sourcing, curating, and analyzing pre-training data for quality and performance.
You’ll work with automated pipelines and human-in-the-loop processes, contributing both scientific insight and production-grade code. It’s ideal for someone who enjoys working at the intersection of data, machine learning, and systems, and who’s excited by the challenge of shaping frontier AI.
This role blends fundamental research and practical engineering, as we do not distinguish between the two roles internally. You will be expected to write high-performance code and read technical reports....
Browse similar: AI jobs · AI startup jobs · Startup jobs · Remote jobs · San Francisco Bay Area