# Site Reliability Engineer, Production at Thinking Machines

- Company: Thinking Machines
- What the company does: Connectionism: Research Blog by Thinking Machines Lab. Backed by a16z, Accel and GV.
- Company website: https://thinkingmachines.ai/
- Type: Startups (AI role)
- Level: Mid level
- Location: San Francisco
- Work setup: On-site
- Pay: $300K to $350K base salary per year (USD)
- Posted: 2026-08-31
- Apply by: 2026-10-15
- Apply: https://jobs.ashbyhq.com/thinkingmachines/c4cddfec-33de-4f67-9801-1308c2151952
- Page: https://www.1752.vc/careers/jobs/thinking-machines-site-reliability-engineer-production/

## About the role

We're looking for a Site Reliability Engineer (SRE) to drive the reliability of Tinker end-to-end. You'll work alongside the engineers building the platform and research teams to make every layer of the system more robust and resilient. Define and own end-to-end reliability, from CI/CD flows to production observability and incident response.

## What they're looking for

- Bachelor's degree or equivalent experience in computer science, engineering, or similar
- Experience in distributed systems, cloud infrastructure, or site reliability engineering
- Proficiency writing software to solve reliability problems, including building tooling and automation
- Experience with production incident response, postmortems, and systematic reliability improvement
- Strong communication skills and track record of coordination across engineering and research teams
- Deep experience operating production cloud services at scale (e.g., public cloud platforms, internal cloud services)

Tags: Core Engineering
