Cerebras powers the world's fastest AI inference on the biggest wafer chip. Cerebras CS-4 delivers up to 30x faster inference than GPUs.
About the role
As a Network Systems Architect, you will define the scale out, and particularly scale-up network architecture for current and future Cerebras platforms, including proprietary accelerator interconnects, protocols, and switching. Requirements will not arrive as a finished bandwidth and latency specification.
What they're looking for
- Relevant experience may include one or more of the following
- • Proprietary accelerator interconnects or scale-up technologies such as NVLink, xGMI, TPU ICI, Xe Link, or UALink-class systems
- • Switch ASIC, NIC or DPU, accelerator I/O, transport offload, collective acceleration, coherent memory, or custom-fabric work
- • FPGA architecture or mapping experience, or work with other placement-sensitive systems where topology materially affects communication
- • AI or HPC communication stacks such as NCCL, RCCL, MPI, SHMEM, or proprietary collective libraries
- • RoCE, InfiniBand, PCIe, CXL, or other relevant scale-out and I/O technologies
More about this role
Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services.
This order of magnitude increase in speed is transforming the user experience of AI applications, unlocking real-time iteration and increasing intelligence via additional agentic computation.
Cerebras works with the leading model labs, global enterprises, and cutting-edge AI-native startups. OpenAI recently announced a multi-year partnership with Cerebras, to deploy 750 megawatts of scale, transforming key workloads with ultra high-speed inference.
As a Network Systems Architect, you will define the scale out, and particularly scale-up network architecture for current and future Cerebras platforms, including proprietary accelerator interconnects, protocols, and switching. Requirements will not arrive as a finished bandwidth and latency specification. Working with application, compiler, runtime, and systems teams, you will study communication patterns, workload partitioning and placement, data and memory movement, synchronization, locality,...
Browse similar: AI jobs · AI startup jobs · Startup jobs · Remote jobs