Connectionism: Research Blog by Thinking Machines Lab. Backed by a16z, Accel and GV.
About the role
As Product Manager for Deployment, you will own how Thinking Machines' models and fine-tuned checkpoints go from training into production use. You will shape the path from a trained model to a served, reliable, cost-effective endpoint — covering inference infrastructure, serving APIs, latency and throughput tradeoffs, scaling behavior, observability, and the workflows researchers and external users rely on to deploy their work with confidence.
What they're looking for
- Experience owning a production ML serving, infrastructure, or deployment product, with direct involvement in reliability, scaling, or performance decisions
- Track record working at engineering depth with production systems — comfortable discussing latency, throughput, autoscaling, rollback, or incident response in specifics
- Experience taking a technical product from early usage through to reliable, scaled production use
- Background as an engineer or technical founder before moving into product leadership
- Experience with ML inference infrastructure specifically (model serving frameworks, GPU scheduling, batching, quantization tradeoffs, or similar)
- Experience operating in a startup, lab, or new product area where the deployment model and roadmap weren't handed to you
More about this role
The mission of Thinking Machines is to build AI that extends human will and judgment. We are training frontier models with Inkling, developing Tinker to let people make models their own, and crafting interfaces that broaden human-AI communication. We believe the future worth building is human, and we're hiring people who want to build it.
As Product Manager for Deployment, you will own how Thinking Machines' models and fine-tuned checkpoints go from training into production use. You will shape the path from a trained model to a served, reliable, cost-effective endpoint — covering inference infrastructure, serving APIs, latency and throughput tradeoffs, scaling behavior, observability, and the workflows researchers and external users rely on to deploy their work with confidence.
This is not a mature MLOps role at an established platform. Deployment at Thinking Machines is still being defined: what "production-ready" means for a fine-tuned model, which serving paths we support, how much control users get over performance and cost tradeoffs, and how we scale reliably as usage grows. You will work from infrastructure capability through to a deployment experience that is fast,...
Browse similar: AI jobs · AI startup jobs · Startup jobs · Remote jobs · San Francisco Bay Area