# Site Reliability Engineer - AI Accelerator Infrastructure - Contract at d-Matrix

- Company: d-Matrix
- What the company does: d-Matrix is redefining AI inference with memory-centric compute built for ultra-low latency, greater efficiency, and scalable AI infrastructure.
- Company website: https://www.d-matrix.ai
- Type: Startups (AI role)
- Level: Mid level
- Location: Santa Clara
- Work setup: Remote
- Posted: 2026-07-14
- Apply by: 2026-10-08
- Apply: https://jobs.ashbyhq.com/d-Matrix/3d4fbebe-c261-47ba-b809-18fc9d25cfcc
- Page: https://www.1752.vc/careers/jobs/d-matrix-site-reliability-engineer-ai-accelerator-infrastructure-contract/

## About the role

We value humility and believe in direct communication. Our team is inclusive , and our differing perspectives allow for better solutions. We are seeking individuals passionate about tackling challenges and are driven by execution. Ready to come find your playground? Together , we can help shape the endless possibilities of AI.

## What they're looking for

- Bachelor's or Master's in Computer Science, Electrical Engineering, or a related field (or equivalent experience), 5+ years in SRE, infrastructure engineering, or systems administration
- Strong Linux systems knowledge: networking, storage, systemd, package management, kernel parameters, and performance diagnostics
- Hands-on experience with colocation or on-premises server infrastructure — physical hardware, rack networking, and bare-metal provisioning
- IaC experience with Terraform and/or Ansible — writing and maintaining production configurations, not just running existing playbooks
- Kubernetes operational experience: cluster troubleshooting, workload management, storage, and networking
- Prometheus + Grafana or DataDog: building dashboards, writing alert rules, and understanding signal quality

Tags: G&A
