About the role
ClusterMAX™ is the industry-standard GPU cloud rating system — 95% market coverage by volume, 84 providers rated, 209 tracked, and 140+ customer interviews behind each release. We recently published our cluster TCO and goodput framework in How Much Do GPU Clusters Really Cost? , showing that goodput expense alone swings 6–21% of total cluster TCO depending on fault-tolerance approach. We are now testing providers for ClusterMAX 3.0 with expanded benchmarks, security requirements, and analysis.
What they're looking for
- Hands-on experience operating GPU clusters: Slurm and/or Kubernetes, InfiniBand/RoCE fabrics, distributed storage
- Strong Python and shell scripting, comfort building benchmark harnesses and CI pipelines
- Understanding of distributed training/inference failure modes — node failures, checkpointing, blast radius, recovery
- Experience similar to our Technical Consultant profile is a plus: due diligence, TCO analysis, and client-facing technical communication
- Security mindset: multi-tenant isolation, bare-metal vs virtualized trade-offs
More about this role
Work Setting: In-office/Remote
Work Location: United States, New York, Mexico, San Francisco, Canada
Work Hours: Office hours
Find out more here: https://semianalysis.com
About SemiAnalysis
SemiAnalysis is an independent research and analysis firm specializing in the Semiconductor and AI industries. Our in-depth coverage spans the entire supply chain, from semiconductor fabrication processes to state-of-the-art AI Models, CUDA kernels, and GPU cloud infrastructure. We are recognized as the leading authority on AI infrastructure, with the highest concentration of industry experts within one team, and a deep-rooted passion for delving into the intricacies.
We’re a global team of over 20 analysts & engineers, each with extensive networks across the semiconductor supply chain and AI ecosystem, publishing industry‑shaping articles while participating in 40+ conferences annually.
Our newsletter reaches more than 200 000 subscribers worldwide, including senior management and C‑suite leaders at the leading semiconductor and AI companies.
InferenceMAX: The world first open inference benchmark that continuous benchmarks performance of popular frontier models
MI300X vs H100 vs H200 Training...
Browse similar: Startup jobs · San Francisco Bay Area