# Site Reliability Engineer, Post Training at Thinking Machines

- Company: Thinking Machines
- What the company does: Connectionism: Research Blog by Thinking Machines Lab. Backed by a16z, Accel and GV.
- Company website: https://thinkingmachines.ai/
- Type: Startups (AI role)
- Level: Mid level
- Location: San Francisco
- Work setup: On-site
- Pay: $300K to $350K base salary per year (USD)
- Posted: 2026-08-31
- Apply by: 2026-10-15
- Apply: https://jobs.ashbyhq.com/thinkingmachines/a9469410-04c7-4e6a-b8b4-64c15933a2bf
- Page: https://www.1752.vc/careers/jobs/thinking-machines-site-reliability-engineer-post-training/

## About the role

We're hiring a Site Reliability Engineer (SRE) to keep our post-training and reinforcement learning (RL) systems fast, reliable, and easy for researchers to iterate on. Think of this as a production engineering or site reliability role built around model training: you'll own the health of the training runs, clusters, and pipelines that power post-training and RL at Thinking Machines.

## What they're looking for

- 4+ years of experience as a production engineer, site reliability engineer, or infrastructure engineer operating large-scale distributed systems in production
- Track record debugging complex failures across distributed systems — networking, hardware, kernel, or scheduler issues
- Strong software engineering skills in Python and/or Go/C++, with the judgment to know when to script a fix versus build a system
- Solid grounding in Linux systems internals and networking fundamentals
- Comfortable owning production systems, including participating in on-call rotations
- Experience operating GPU or TPU training clusters at scale

Tags: Core Engineering
