# Member of Technical Staff — Compute Cluster at Causal

- Company: Causal
- What the company does: Backed by Accel.
- Company website: https://www.causal.app/
- Type: Startups
- Level: Senior
- Location: San Francisco
- Work setup: On-site
- Posted: 2026-07-19
- Apply by: 2026-10-08
- Apply: https://jobs.ashbyhq.com/causal/ab4edc29-4427-4e02-88c4-bc73adf65df8
- Page: https://www.1752.vc/careers/jobs/causal-member-of-technical-staff-compute-cluster/

## About the role

Design, deploy, and operate large distributed GPU clusters end to end: provisioning, imaging, upgrades, and capacity planning Extend scheduling and orchestration systems (e.g. Kubernetes, Slurm) for topology-aware placement, preemption, quotas, and multi-tenancy across training and inference workloads

## What they're looking for

- We value a relentless approach to problem-solving, rapid execution, and the ability to quickly learn in unfamiliar domains
- Experience operating large-scale GPU clusters and container orchestration frameworks (e.g. Kubernetes, Slurm, Docker)
- Strong systems background: Linux, networking, storage, infrastructure-as-code
- Knowledge of cloud platforms (GCP, AWS, or Azure) and their ML/AI service offerings
- Understanding of monitoring, logging, observability, and version control best practices for ML systems
- Familiarity with CUDA/NCCL and performance profiling for distributed workloads

Tags: Infrastructure
