# Director, Site Reliability Engineering - AI Accelerator Infrastructure - Contract at d-Matrix

- Company: d-Matrix
- What the company does: d-Matrix is redefining AI inference with memory-centric compute built for ultra-low latency, greater efficiency, and scalable AI infrastructure.
- Company website: https://www.d-matrix.ai
- Type: Startups (AI role)
- Level: Principal and up
- Location: Santa Clara
- Work setup: Remote
- Pay: $195K to $285K base salary per year (USD)
- Posted: 2026-07-13
- Apply by: 2026-10-08
- Apply: https://jobs.ashbyhq.com/d-Matrix/36e787a8-966f-46f8-87f3-7572384b1839
- Page: https://www.1752.vc/careers/jobs/d-matrix-director-site-reliability-engineering-ai-accelerator-infrastructure-con/

## About the role

You will build and lead d-Matrix’s site reliability engineering function from the ground up—owning the infrastructure that development, validation, and customer-facing deployments run on. This spans colocation facilities, on-premises lab clusters, cloud environments (AWS, Azure, and GCP), and the platform services customers use to collaborate with d-Matrix on hardware and software deployments.

## What they're looking for

- Bachelor’s or Master’s in Computer Science, Electrical Engineering, or a related field, 15+ years in SRE, infrastructure engineering, or production engineering
- 5+ years leading SRE or infrastructure engineering teams — including experience building or significantly rebuilding a function, not just managing a steady-state team
- Demonstrated track record of establishing SRE as a discipline in an organization that lacked it: defining SLOs, creating on-call frameworks, standing up observability, and driving cultural change with engineering teams that came from a reactive ops background
- Proven experience operating colocation and on-premises hardware at scale: server lifecycle, power and cooling awareness, rack-level networking
- IaC fluency: Terraform and Ansible at production scale — module design, remote state, environment isolation, and change governance
- Kubernetes cluster operations: lifecycle management, workload reliability, storage, and RBAC at scale

Tags: G&A
