Startups · AI

Product Manager - Deployment

Thinking Machines · San Francisco · Remote

← All jobs
About Thinking Machines

Connectionism: Research Blog by Thinking Machines Lab. Backed by a16z, Accel and GV.

About the role

As Product Manager for Deployment, you will own how Thinking Machines' models and fine-tuned checkpoints go from training into production use. You will shape the path from a trained model to a served, reliable, cost-effective endpoint — covering inference infrastructure, serving APIs, latency and throughput tradeoffs, scaling behavior, observability, and the workflows researchers and external users rely on to deploy their work with confidence.

What they're looking for

  • Experience owning a production ML serving, infrastructure, or deployment product, with direct involvement in reliability, scaling, or performance decisions
  • Track record working at engineering depth with production systems — comfortable discussing latency, throughput, autoscaling, rollback, or incident response in specifics
  • Experience taking a technical product from early usage through to reliable, scaled production use
  • Background as an engineer or technical founder before moving into product leadership
  • Experience with ML inference infrastructure specifically (model serving frameworks, GPU scheduling, batching, quantization tradeoffs, or similar)
  • Experience operating in a startup, lab, or new product area where the deployment model and roadmap weren't handed to you
More about this role

The mission of Thinking Machines is to build AI that extends human will and judgment. We are training frontier models with Inkling, developing Tinker to let people make models their own, and crafting interfaces that broaden human-AI communication. We believe the future worth building is human, and we're hiring people who want to build it.

As Product Manager for Deployment, you will own how Thinking Machines' models and fine-tuned checkpoints go from training into production use. You will shape the path from a trained model to a served, reliable, cost-effective endpoint — covering inference infrastructure, serving APIs, latency and throughput tradeoffs, scaling behavior, observability, and the workflows researchers and external users rely on to deploy their work with confidence.

This is not a mature MLOps role at an established platform. Deployment at Thinking Machines is still being defined: what "production-ready" means for a fine-tuned model, which serving paths we support, how much control users get over performance and cost tradeoffs, and how we scale reliably as usage grows. You will work from infrastructure capability through to a deployment experience that is fast,...

Read the full posting on Thinking Machines's site ↗

Product Management

Build your edge while you search

Free tools for founders and investors, plus VC Unfiltered, our take on startups, venture and the people who build them.