Cerebras powers the world's fastest AI inference on the biggest wafer chip. Cerebras CS-4 delivers up to 30x faster inference than GPUs.
About the role
Own and evolve performance models and modeling methodologies for next-generation accelerator and system architectures. Build and extend analytical, simulation-based or trace-driven models across workloads, architectural features and product generations.
What they're looking for
- 7+ years of experience in performance analysis, performance modeling or architecture exploration for CPUs, GPUs, AI accelerators or other high-performance computing systems
- Strong understanding of hardware architecture developed through hardware, compiler, kernel, runtime or system-performance work
- Experience developing analytical, simulation-based or trace-driven performance models using Python, C++ or similar environments
- Solid understanding of processor architecture, memory systems, interconnects, parallel execution and hardware resource constraints
- Ability to move between kernel-level behavior and end-to-end application or system performance
- Experience profiling workloads, forming performance hypotheses and validating them with quantitative evidence
More about this role
Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services.
This order of magnitude increase in speed is transforming the user experience of AI applications, unlocking real-time iteration and increasing intelligence via additional agentic computation.
Cerebras works with the leading model labs, global enterprises, and cutting-edge AI-native startups. OpenAI recently announced a multi-year partnership with Cerebras, to deploy 750 megawatts of scale, transforming key workloads with ultra high-speed inference.
Wafer-scale computing creates a distinctive architecture space in which compute placement, memory capacity and bandwidth, communication, kernel execution and system-level behavior must be understood together.
We are looking for a performance architect to guide the evolution of our next-generation AI systems. You will connect real workloads to architectural behavior, identify the bottlenecks that matter, quantify potential improvements and influence hardware and software roadmaps through rigorous...
Browse similar: AI jobs · AI startup jobs · Startup jobs