Connectionism: Research Blog by Thinking Machines Lab. Backed by a16z, Accel and GV.
About the role
Our team scales reinforcement learning for frontier models. Progress in RL is increasingly set by how well it scales: more rollouts, larger models, and training loops that keep large fleets of accelerators doing useful work. We are particularly interested in people working on high-training-compute, long-horizon RL. We believe the biggest gains come from designing the training recipe and the infrastructure together rather than separately, and we are hiring a researcher who wants to own that boundary.
What they're looking for
- Proficiency in Python and familiarity with at least one deep learning framework (e.g., PyTorch, TensorFlow, or JAX). Comfortable with debugging distributed training and writing code that scales
- Bachelor’s degree or equivalent experience in Computer Science, Machine Learning, Physics, Mathematics, or a related discipline with strong theoretical and empirical grounding
- Clarity in communication, an ability to explain complex technical concepts in writing
- Strong research judgment: clean ablations, honest baselines, and clear technical writing
- Preferred qualifications — we encourage you to apply if you meet some but not all of these:
- PhD in Computer Science, Machine Learning, Physics, Mathematics, or a related discipline with strong theoretical and empirical grounding, or, equivalent industry research experience
More about this role
The mission of Thinking Machines is to build AI that extends human will and judgment. We are training frontier models with Inkling, developing Tinker to let people make models their own, and crafting interfaces that broaden human-AI communication. We believe the future worth building is human, and we're hiring people who want to build it.
Our team scales reinforcement learning for frontier models. Progress in RL is increasingly set by how well it scales: more rollouts, larger models, and training loops that keep large fleets of accelerators doing useful work. We are particularly interested in people working on high-training-compute, long-horizon RL. We believe the biggest gains come from designing the training recipe and the infrastructure together rather than separately, and we are hiring a researcher who wants to own that boundary.
A center of gravity for this role is asynchronous RL. Decoupling generation from training changes both the systems design and the learning problem, and doing it well requires a deep understanding of async RL algorithms, design choices, and trade-offs on both the ML and the systems sides. We expect much of the headroom in RL scaling to come from...
Browse similar: AI jobs · AI startup jobs · Startup jobs · Remote jobs · San Francisco Bay Area