Startups · AI

Site Reliability Engineer, Post Training

Thinking Machines · San Francisco · On-site

← All jobs
About Thinking Machines

Connectionism: Research Blog by Thinking Machines Lab. Backed by a16z, Accel and GV.

About the role

We're hiring a Site Reliability Engineer (SRE) to keep our post-training and reinforcement learning (RL) systems fast, reliable, and easy for researchers to iterate on. Think of this as a production engineering or site reliability role built around model training: you'll own the health of the training runs, clusters, and pipelines that power post-training and RL at Thinking Machines.

What they're looking for

  • 4+ years of experience as a production engineer, site reliability engineer, or infrastructure engineer operating large-scale distributed systems in production
  • Track record debugging complex failures across distributed systems — networking, hardware, kernel, or scheduler issues
  • Strong software engineering skills in Python and/or Go/C++, with the judgment to know when to script a fix versus build a system
  • Solid grounding in Linux systems internals and networking fundamentals
  • Comfortable owning production systems, including participating in on-call rotations
  • Experience operating GPU or TPU training clusters at scale
More about this role

The mission of Thinking Machines is to build AI that extends human will and judgment. We are training frontier models with Inkling, developing Tinker to let people make models their own, and crafting interfaces that broaden human-AI communication. We believe the future worth building is human, and we're hiring people who want to build it.

We're hiring a Site Reliability Engineer (SRE) to keep our post-training and reinforcement learning (RL) systems fast, reliable, and easy for researchers to iterate on. Think of this as a production engineering or site reliability role built around model training: you'll own the health of the training runs, clusters, and pipelines that power post-training and RL at Thinking Machines.

You'll work side by side with research teams during active model runs — debugging failures in real time, hardening infrastructure against the next class of problem, and building the tooling and automation that let researchers spend their time on the science instead of babysitting jobs. This role has real ownership: you'll be the person a research team calls when a run stalls at 2am, and the person who makes sure it doesn't happen again.

Own the reliability,...

Read the full posting on Thinking Machines's site ↗

Core Engineering

Build your edge while you search

Free tools for founders and investors, plus VC Unfiltered, our take on startups, venture and the people who build them.