Connectionism: Research Blog by Thinking Machines Lab. Backed by a16z, Accel and GV.
About the role
Mid-training is a step between pre-training and post-training, where we take a base model and train it into the foundation for reasoning. This role owns the late-stage training responsibility that shape what our models are fundamentally capable of, including things like synthetic data strategies, the data mix, quality uplift, context extension and capabilities across coding, math, reasoning, and so on.
What they're looking for
- Proficiency in Python and familiarity with deep learning frameworks (e.g., PyTorch, TensorFlow, or JAX). Comfort debugging distributed training and writing code that scales
- Bachelor’s degree or equivalent experience in Computer Science, Machine Learning, Physics, Mathematics, or a related discipline with strong theoretical and empirical grounding
- Clarity in communication, an ability to explain complex technical concepts in writing
- Preferred qualifications (we encourage you to apply if you meet some but not all of these):
- A strong grasp of probability, statistics, and ML fundamentals. You can look at experimental data and distinguish between real effects, noise, and bugs
- Experience building or owning training datasets for large models, including both synthetic data pipelines and real user data (collection, curation, filtering, or mixture design)
More about this role
The mission of Thinking Machines is to build AI that extends human will and judgment. We are training frontier models with Inkling, developing Tinker to let people make models their own, and crafting interfaces that broaden human-AI communication. We believe the future worth building is human, and we're hiring people who want to build it.
Mid-training is a step between pre-training and post-training, where we take a base model and train it into the foundation for reasoning. This role owns the late-stage training responsibility that shape what our models are fundamentally capable of, including things like synthetic data strategies, the data mix, quality uplift, context extension and capabilities across coding, math, reasoning, and so on.
This role blends fundamental research and practical engineering, as we do not distinguish between the two internally. It's an excellent fit for someone comfortable working across the boundary of pre-training and post-training, and who wants to shape what our models can do at their core.
What You’ll Do
Own the data. Decide what the model needs to see for each capability and each area of knowledge, then source, curate, and synthesize it. Build the...
Browse similar: AI jobs · AI startup jobs · Startup jobs · Remote jobs · San Francisco Bay Area