# Director Site Reliability Engineer, AI Infrastructure at d-Matrix

- Company: d-Matrix
- What the company does: d-Matrix is redefining AI inference with memory-centric compute built for ultra-low latency, greater efficiency, and scalable AI infrastructure.
- Company website: https://www.d-matrix.ai
- Type: Startups (AI role)
- Level: Principal and up
- Location: Santa Clara
- Work setup: Remote
- Pay: $195K to $285K base salary per year (USD)
- Posted: 2026-09-26
- Apply by: 2026-11-10
- Apply: https://jobs.ashbyhq.com/d-Matrix/505b7fa6-666a-44d5-9b23-88b7eec20502
- Page: https://www.1752.vc/careers/jobs/d-matrix-director-site-reliability-engineer-ai-infrastructure/

## About the role

d-Matrix designs and manufactures purpose-built AI inference silicon, and the infrastructure underpinning our engineering organization must be as reliable and scalable as the chips we build. This role builds and leads d-Matrix's Site Reliability Engineering function from the ground up, owning the infrastructure that development, validation, and customer-facing deployments run on — spanning colocation facilities, on-premises lab clusters, cloud environments (AWS, Azure, GCP), and customer-facing platform services.

## What they're looking for

- Experience operating customer-facing infrastructure or platform services, with reliability expectations beyond internal tooling
- Knowledge of high-speed interconnect fabrics (InfiniBand, RoCE, or NVLink) — setup, troubleshooting, and performance tuning
- HPC job scheduler experience (Slurm, LSF, or equivalent) — setup, tuning, and integration with infrastructure automation
- Multi-cloud hybrid operations across AWS, Azure, and GCP alongside on-prem/colo, with unified observability and IaC
- FinOps expertise: cloud spend attribution, TCO modeling across cloud vs. on-prem vs. colo, and translating cost data into workload placement recommendations
- ITIL knowledge or an equivalent structured incident/problem/change management framework

Tags: G&A
