Power AI workloads with AMD Instinct™ GPUs. Scale models faster with high-performance, dedicated cloud compute.
About the role
We are building our level 2 technical support team from the ground up, and we are looking for a Technical Support Engineer to join it.
What they're looking for
- 3+ years in Linux systems administration, site reliability engineering, infrastructure operations, or a technical support engineering role with real diagnostic ownership
- Strong Linux troubleshooting depth: system, networking, file systems, storage, process and resource investigation, log analysis, and comfort at the kernel boundary when applications go quiet
- Working knowledge of Kubernetes, reading cluster and pod state, understanding scheduling and placement, and diagnosing why a workload will not run
- Experience with a batch scheduler in a shared compute environment, ideally Slurm, including job submission, queue behavior, and node state
- Scripting ability in Python or Bash, sufficient to automate diagnosis
- Experience working tickets in a structured incident and tracking system such as JIRA, PagerDuty, or equivalent
More about this role
Our mission is simple: deliver seamless, secure, reliable, and resilient AI compute at scale. We've built a versatile cloud platform that eliminates infrastructure barriers, empowering builders to focus on innovation instead of fighting their stack. Because breakthrough AI should move at the speed of ideas, not infrastructure.
We are building our level 2 technical support team from the ground up, and we are looking for a Technical Support Engineer to join it.
Our customers are not filing tickets about forgotten passwords; they are AI companies running massive GPU training jobs and production inference at scale, and when they contact us, something real is wrong. A node is reporting 7 GPUs instead of 8. A training run that has been going for 4 days is suddenly crawling, and nobody can say why. A Slurm partition is draining, and the queue is backing up. Our global operations center catches and handles what the runbooks cover, around the clock, everything past that comes to you.
This is a hands-on diagnostic role for someone who genuinely enjoys the hunt. You will work Linux systems at depth, live inside Kubernetes and Slurm, read logs and metrics until the story makes sense, and talk...
Browse similar: Startup jobs