Startups · AI

Director, Site Reliability Engineering - AI Accelerator Infrastructure - Contract

d-Matrix · Santa Clara · Remote

← All jobs
About d-Matrix

d-Matrix is redefining AI inference with memory-centric compute built for ultra-low latency, greater efficiency, and scalable AI infrastructure.

About the role

You will build and lead d-Matrix’s site reliability engineering function from the ground up—owning the infrastructure that development, validation, and customer-facing deployments run on. This spans colocation facilities, on-premises lab clusters, cloud environments (AWS, Azure, and GCP), and the platform services customers use to collaborate with d-Matrix on hardware and software deployments.

What they're looking for

  • Bachelor’s or Master’s in Computer Science, Electrical Engineering, or a related field, 15+ years in SRE, infrastructure engineering, or production engineering
  • 5+ years leading SRE or infrastructure engineering teams — including experience building or significantly rebuilding a function, not just managing a steady-state team
  • Demonstrated track record of establishing SRE as a discipline in an organization that lacked it: defining SLOs, creating on-call frameworks, standing up observability, and driving cultural change with engineering teams that came from a reactive ops background
  • Proven experience operating colocation and on-premises hardware at scale: server lifecycle, power and cooling awareness, rack-level networking
  • IaC fluency: Terraform and Ansible at production scale — module design, remote state, environment isolation, and change governance
  • Kubernetes cluster operations: lifecycle management, workload reliability, storage, and RBAC at scale
More about this role

At d-Matrix , we are focused on unleashing the potential of generative AI to power the transformation of technology. We are at the forefront of software and hardware innovation, pushing the boundaries of what is possible. Our culture is one of respect and collaboration.

We value humility and believe in direct communication. Our team is inclusive , and our differing perspectives allow for better solutions. We are seeking individuals passionate about tackling challenges and are driven by execution. Ready to come find your playground? Together , we can help shape the endless possibilities of AI.

Director, Site Reliability Engineering - AI Accelerator Infrastructure - Contract

About d-Matrix

d-Matrix designs and manufactures purpose-built AI inference silicon. Our engineering organization spans silicon, software, hardware, QA, and research, and the infrastructure that underpins all of it must be as reliable and scalable as the chips we build.

We compete directly with Nvidia for engineering talent and hold ourselves to the same bar—the infrastructure organization is no exception. The SRE team owns the physical and virtual infrastructure layer that the entire company — and our customers...

Read the full posting on d-Matrix's site ↗

G&A

Build your edge while you search

Free tools for founders and investors, plus VC Unfiltered, our take on startups, venture and the people who build them.