Startups · AI

Sr Staff Site Reliability Engineer, AI Infrastructure

d-Matrix · Santa Clara · Remote

← All jobs
About d-Matrix

d-Matrix is redefining AI inference with memory-centric compute built for ultra-low latency, greater efficiency, and scalable AI infrastructure.

About the role

d-Matrix's SRE team owns the infrastructure layer that every engineering team and customer depends on — colocation facilities, on-premises GPU clusters, cloud environments, and the platform services used to deploy and validate d-Matrix hardware and software. This role is a core member of that team, responsible for reliability, automation, and observability across colo, on-premises lab, and cloud environments. Found on 1752vc Careers, the job board for startup and VC roles.

What they're looking for

  • Experience operating customer-facing infrastructure or platform services with external reliability expectations
  • Cloud infrastructure operations across AWS, Azure, or GCP, including hybrid environments spanning cloud and on-prem
  • Experience deploying and operating AI-driven infrastructure tools — AIOps platforms, intelligent alerting, anomaly detection, or LLM-assisted diagnostics — in production
  • HPC job scheduler experience: Slurm, LSF, or equivalent
  • Knowledge of high-speed interconnect fabrics: InfiniBand, RoCE, or NVLink
  • Experience with large-scale infrastructure automation — host lifecycle management, fleet auto-healing, or AIOps-driven operations — building tooling that reduces manual intervention, not just running it
More about this role

At d-Matrix , we are focused on unleashing the potential of generative AI to power the transformation of technology. We are at the forefront of software and hardware innovation, pushing the boundaries of what is possible. Our culture is one of respect and collaboration.

We value humility and believe in direct communication. Our team is inclusive , and our differing perspectives allow for better solutions. We are seeking individuals passionate about tackling challenges and are driven by execution. Ready to come find your playground? Together , we can help shape the endless possibilities of AI.

d-Matrix's SRE team owns the infrastructure layer that every engineering team and customer depends on — colocation facilities, on-premises GPU clusters, cloud environments, and the platform services used to deploy and validate d-Matrix hardware and software. This role is a core member of that team, responsible for reliability, automation, and observability across colo, on-premises lab, and cloud environments. You will own systems end-to-end, from provisioning through live incident response, partnering with hardware and software teams on CI/CD, QA, and HPC workloads for silicon development, as...

Read the full posting on d-Matrix's site ↗

G&A

Build your edge while you search

Free tools for founders and investors, plus VC Unfiltered, our take on startups, venture and the people who build them.