# Sr Staff Site Reliability Engineer, AI Infrastructure at d-Matrix

- Company: d-Matrix
- What the company does: d-Matrix is redefining AI inference with memory-centric compute built for ultra-low latency, greater efficiency, and scalable AI infrastructure.
- Company website: https://www.d-matrix.ai
- Type: Startups (AI role)
- Level: Senior
- Location: Santa Clara
- Work setup: Remote
- Pay: $175K to $265K base salary per year (USD)
- Posted: 2026-09-29
- Apply by: 2026-11-13
- Apply: https://jobs.ashbyhq.com/d-Matrix/e4cd2043-5639-44b3-aed1-bbb0ad987709?utm_source=1752vc&utm_medium=careers
- Page: https://www.1752.vc/careers/jobs/d-matrix-sr-staff-site-reliability-engineer-ai-infrastructure/

## About the role

d-Matrix's SRE team owns the infrastructure layer that every engineering team and customer depends on — colocation facilities, on-premises GPU clusters, cloud environments, and the platform services used to deploy and validate d-Matrix hardware and software. This role is a core member of that team, responsible for reliability, automation, and observability across colo, on-premises lab, and cloud environments. Found on 1752vc Careers, the job board for startup and VC roles.

## What they're looking for

- Experience operating customer-facing infrastructure or platform services with external reliability expectations
- Cloud infrastructure operations across AWS, Azure, or GCP, including hybrid environments spanning cloud and on-prem
- Experience deploying and operating AI-driven infrastructure tools — AIOps platforms, intelligent alerting, anomaly detection, or LLM-assisted diagnostics — in production
- HPC job scheduler experience: Slurm, LSF, or equivalent
- Knowledge of high-speed interconnect fabrics: InfiniBand, RoCE, or NVLink
- Experience with large-scale infrastructure automation — host lifecycle management, fleet auto-healing, or AIOps-driven operations — building tooling that reduces manual intervention, not just running it

Tags: G&A

Source: 1752vc Careers, https://www.1752.vc/careers/jobs/d-matrix-sr-staff-site-reliability-engineer-ai-infrastructure/
