Startups · AI

Cluster Operations Software Engineer

Cerebras Systems · Sunnyvale, CA · Remote

← All jobs
About Cerebras Systems

Cerebras powers the world's fastest AI inference on the biggest wafer chip. Cerebras CS-4 delivers up to 30x faster inference than GPUs.

About the role

We are seeking a highly skilled and experienced AI Cluster Operations Engineer to manage and operate our cutting-edge machine learning compute clusters. These clusters would provide the candidate with an opportunity to work with the world's largest computer chip, the Wafer-Scale Engine (WSE), and the systems that harness its unparalleled power.

What they're looking for

  • 6-8 years of relevant experience in managing and operating complex compute infrastructure, preferably in the context of machine learning or high-performance computing
  • Proficient in Python and Go, with experience building operational platforms, workflow automation systems, and reliability tooling for large-scale infrastructure environments
  • Experience and Expertise in distributed systems is a must
  • Deep understanding of Linux-based compute systems and command-line tools
  • Extensive knowledge of Docker containers and container orchestration platforms like k8s
  • Proven ability to troubleshoot and resolve complex technical issues in a timely and efficient manner
More about this role

Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services.

This order of magnitude increase in speed is transforming the user experience of AI applications, unlocking real-time iteration and increasing intelligence via additional agentic computation.

Cerebras works with the leading model labs, global enterprises, and cutting-edge AI-native startups. OpenAI recently announced a multi-year partnership with Cerebras, to deploy 750 megawatts of scale, transforming key workloads with ultra high-speed inference.

We are seeking a highly skilled and experienced AI Cluster Operations Engineer to manage and operate our cutting-edge machine learning compute clusters. These clusters would provide the candidate with an opportunity to work with the world's largest computer chip, the Wafer-Scale Engine (WSE), and the systems that harness its unparalleled power.

You will play a critical role in ensuring the health, performance, and availability of our infrastructure, maximizing compute capacity, and supporting...

Read the full posting on Cerebras Systems's site ↗

Software Engineering

Build your edge while you search

Free tools for founders and investors, plus VC Unfiltered, our take on startups, venture and the people who build them.