Cerebras powers the world's fastest AI inference on the biggest wafer chip. Cerebras CS-4 delivers up to 30x faster inference than GPUs.
About the role
We are seeking a highly skilled and experienced AI Cluster Operations Engineer to manage and operate our cutting-edge machine learning compute clusters. These clusters would provide the candidate with an opportunity to work with the world's largest computer chip, the Wafer-Scale Engine (WSE), and the systems that harness its unparalleled power.
What they're looking for
- 6-8 years of relevant experience in managing and operating complex compute infrastructure, preferably in the context of machine learning or high-performance computing
- Proficient in Python and Go, with experience building operational platforms, workflow automation systems, and reliability tooling for large-scale infrastructure environments
- Experience and Expertise in distributed systems is a must
- Deep understanding of Linux-based compute systems and command-line tools
- Extensive knowledge of Docker containers and container orchestration platforms like k8s
- Proven ability to troubleshoot and resolve complex technical issues in a timely and efficient manner
More about this role
Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services.
This order of magnitude increase in speed is transforming the user experience of AI applications, unlocking real-time iteration and increasing intelligence via additional agentic computation.
Cerebras works with the leading model labs, global enterprises, and cutting-edge AI-native startups. OpenAI recently announced a multi-year partnership with Cerebras, to deploy 750 megawatts of scale, transforming key workloads with ultra high-speed inference.
We are seeking a highly skilled and experienced AI Cluster Operations Engineer to manage and operate our cutting-edge machine learning compute clusters. These clusters would provide the candidate with an opportunity to work with the world's largest computer chip, the Wafer-Scale Engine (WSE), and the systems that harness its unparalleled power.
You will play a critical role in ensuring the health, performance, and availability of our infrastructure, maximizing compute capacity, and supporting...
Browse similar: AI jobs · AI startup jobs · Startup jobs · Remote jobs