d-Matrix is redefining AI inference with memory-centric compute built for ultra-low latency, greater efficiency, and scalable AI infrastructure.
About the role
We value humility and believe in direct communication. Our team is inclusive , and our differing perspectives allow for better solutions. We are seeking individuals passionate about tackling challenges and are driven by execution. Ready to come find your playground? Together , we can help shape the endless possibilities of AI.
What they're looking for
- Bachelor's or Master's in Computer Science, Electrical Engineering, or a related field (or equivalent experience), 5+ years in SRE, infrastructure engineering, or systems administration
- Strong Linux systems knowledge: networking, storage, systemd, package management, kernel parameters, and performance diagnostics
- Hands-on experience with colocation or on-premises server infrastructure — physical hardware, rack networking, and bare-metal provisioning
- IaC experience with Terraform and/or Ansible — writing and maintaining production configurations, not just running existing playbooks
- Kubernetes operational experience: cluster troubleshooting, workload management, storage, and networking
- Prometheus + Grafana or DataDog: building dashboards, writing alert rules, and understanding signal quality
More about this role
At d-Matrix , we are focused on unleashing the potential of generative AI to power the transformation of technology. We are at the forefront of software and hardware innovation, pushing the boundaries of what is possible. Our culture is one of respect and collaboration.
We value humility and believe in direct communication. Our team is inclusive , and our differing perspectives allow for better solutions. We are seeking individuals passionate about tackling challenges and are driven by execution. Ready to come find your playground? Together , we can help shape the endless possibilities of AI.
Site Reliability Engineer - AI Accelerator Infrastructure - Contract
About d-Matrix
d-Matrix designs and manufactures purpose-built AI inference silicon. Our SRE team owns the infrastructure layer that every engineering team — and our customers — depends on. That means colocation facilities, on-premises GPU clusters, cloud environments, and the platform services customers use when deploying and validating d-Matrix hardware and software.
This is a 6-month contract with potential conversion to full-time. It is a hands-on, high-ownership role. You will build, operate, and automate real...
Browse similar: AI jobs · AI startup jobs · Startup jobs · Remote jobs