d-Matrix is redefining AI inference with memory-centric compute built for ultra-low latency, greater efficiency, and scalable AI infrastructure.
About the role
d-Matrix designs and manufactures purpose-built AI inference silicon, and the infrastructure underpinning our engineering organization must be as reliable and scalable as the chips we build. This role builds and leads d-Matrix's Site Reliability Engineering function from the ground up, owning the infrastructure that development, validation, and customer-facing deployments run on — spanning colocation facilities, on-premises lab clusters, cloud environments (AWS, Azure, GCP), and customer-facing platform services.
What they're looking for
- Experience operating customer-facing infrastructure or platform services, with reliability expectations beyond internal tooling
- Knowledge of high-speed interconnect fabrics (InfiniBand, RoCE, or NVLink) — setup, troubleshooting, and performance tuning
- HPC job scheduler experience (Slurm, LSF, or equivalent) — setup, tuning, and integration with infrastructure automation
- Multi-cloud hybrid operations across AWS, Azure, and GCP alongside on-prem/colo, with unified observability and IaC
- FinOps expertise: cloud spend attribution, TCO modeling across cloud vs. on-prem vs. colo, and translating cost data into workload placement recommendations
- ITIL knowledge or an equivalent structured incident/problem/change management framework
More about this role
At d-Matrix , we are focused on unleashing the potential of generative AI to power the transformation of technology. We are at the forefront of software and hardware innovation, pushing the boundaries of what is possible. Our culture is one of respect and collaboration.
We value humility and believe in direct communication. Our team is inclusive , and our differing perspectives allow for better solutions. We are seeking individuals passionate about tackling challenges and are driven by execution. Ready to come find your playground? Together , we can help shape the endless possibilities of AI.
d-Matrix designs and manufactures purpose-built AI inference silicon, and the infrastructure underpinning our engineering organization must be as reliable and scalable as the chips we build. This role builds and leads d-Matrix's Site Reliability Engineering function from the ground up, owning the infrastructure that development, validation, and customer-facing deployments run on — spanning colocation facilities, on-premises lab clusters, cloud environments (AWS, Azure, GCP), and customer-facing platform services. You will hire and grow the team, set technical direction, own SLOs for critical...
Browse similar: AI jobs · AI startup jobs · Startup jobs · Remote jobs