Startups

Site Reliability Engineer

Instrumental · Palo Alto, United States · On-site

← All jobs
About Instrumental

Backed by First Round Capital.

About the role

As a Site Reliability Engineer , you’ll operate, improve, and scale our AWS-based SaaS platform. You’ll combine hands-on production operations with engineering, focusing on reliability, automation, observability, and operational excellence. You’ll participate in a bi-weekly on-call rotation, but the goal isn’t simply to keep systems running—it’s to continuously engineer away the operational complexity that comes with scaling our platform and customer base.

What they're looking for

  • 3–4 years of experience in Site Reliability Engineering, DevOps, Cloud Operations, Platform Engineering, or Systems Engineering supporting production SaaS environments
  • Strong hands-on experience with AWS, including EC2, VPC, IAM, RDS, ECS, and S3
  • Experience managing infrastructure using Terraform or other Infrastructure as Code technologies
  • Experience designing and supporting CI/CD pipelines using GitHub Actions, Jenkins, GitLab CI/CD, or similar platforms
  • Strong experience with monitoring and observability tools, preferably Datadog, including dashboards, alerting, logging, and APM
  • Experience with Docker and Kubernetes
More about this role

Instrumental builds the manufacturing acceleration platform behind the world’s most complex electronics. We capture digital exhaust and engineering context from assembly lines—images, test logs, BOM data, performance, repair cycles—and our AI engines identify insights that are difficult or impossible for human engineers to find. We accelerate the companies building the AI era by improving manufacturing yield, throughput, and ramp. NVIDIA, Meta, Cisco, and their manufacturing partners rely on Instrumental to accelerate new product introduction and production.

The Instrumental platform collects, intelligently transforms, and contextually presents manufacturing data to technical end-users, enabling them to optimize their manufacturing process in real-time. Our core technology is proprietary ML algorithms, packaged in an accessible, user-centric user interface—we believe we must have both the best technology and the best access to that technology to win.

As a Site Reliability Engineer , you’ll operate, improve, and scale our AWS-based SaaS platform. You’ll combine hands-on production operations with engineering, focusing on reliability, automation, observability, and operational...

Read the full posting on Instrumental's site ↗

Build your edge while you search

Free tools for founders and investors, plus VC Unfiltered, our take on startups, venture and the people who build them.