# HPC Infrastructure Engineer - GPU Clusters at ElevenLabs

- Company: ElevenLabs
- What the company does: Create lifelike speech with our AI voice generator and voice agents platform. Access 5,000+ voices in 70+ languages with secure APIs and SDKs. Backed by Lightspeed, NEA and Sequoia.
- Company website: https://elevenlabs.io/
- Type: Startups (AI role)
- Level: Mid level
- Location: United States
- Work setup: Remote
- Posted: 2026-09-03
- Apply by: 2026-10-18
- Apply: https://jobs.ashbyhq.com/elevenlabs/120da2b3-d88b-4e3c-9b89-d19ff73db9d9
- Page: https://www.1752.vc/careers/jobs/elevenlabs-hpc-infrastructure-engineer-gpu-clusters/

## About the role

Every model we train runs on infrastructure this role owns. We operate NVIDIA GPU clusters across bare metal and rented capacity, and we're looking for an engineer to join our small research infrastructure team and make that compute fast, reliable, and boring - in the best sense. When the clusters just work, research moves faster. Your impact is measured directly in training throughput and researcher velocity.

## What they're looking for

- Have run large-scale Linux server or GPU environments in production and enjoy both building and operating
- Know the NVIDIA stack well — drivers, CUDA, NCCL, DCGM — or have deep systems experience and learn hardware stacks fast
- Are comfortable with bare-metal environments, server hardware, and high-speed networking
- Write solid automation in Python and/or Bash, with IaC tools like Ansible or Terraform
- Are happy digging into noisy data (metrics, logs, PromQL) to find what's actually wrong
- Like owning real scope end to end and being the person others rely on

Tags: Research & Research Engineering
