Startups · AI

Member of Technical Staff - Infrastructure

Gimletlabs · San Francisco, CA · On-site

← All jobs
About Gimletlabs

Backed by Menlo and Sapphire.

About the role

As an Infrastructure Platform Engineer, you will build the systems that turn heterogeneous accelerator hardware into reliable production infrastructure for Gimlet's AI cloud. Gimlet's fleet spans hardware with different architectures, software stacks, operational characteristics, and failure modes. Your work will determine how new hardware is brought online, how clusters are provisioned and operated, and how production inference systems remain reliable as the fleet scales.

What they're looking for

  • Experience in infrastructure, cluster engineering, platform engineering, SRE, or HPC
  • Strong Linux systems knowledge and production debugging experience
  • Experience operating Kubernetes, Slurm, Nomad, or similar orchestration systems
  • Experience automating infrastructure with Python, Go, Terraform, Ansible, or similar tools
  • Experience with GPU or accelerator infrastructure, including drivers, firmware, or CUDA/ROCm
  • The ability to build systems that are observable, recoverable, and reliable in production
More about this role

Gimlet is building the first multi-silicon neocloud designed for fast, efficient AI inference.

We combine large-scale compute infrastructure with an execution platform that partitions AI workloads and maps each stage to the hardware best suited to run it.

We work with foundation labs, hyperscalers, and AI-native companies, giving our team access to technical problems spanning frontier models, production infrastructure, and emerging hardware.

As an Infrastructure Platform Engineer, you will build the systems that turn heterogeneous accelerator hardware into reliable production infrastructure for Gimlet's AI cloud.

Gimlet's fleet spans hardware with different architectures, software stacks, operational characteristics, and failure modes. Your work will determine how new hardware is brought online, how clusters are provisioned and operated, and how production inference systems remain reliable as the fleet scales.

You will work across bare metal, Linux, Kubernetes, cluster scheduling, observability, and automation. You will build systems that abstract differences between accelerator architectures, make new hardware production-ready, and improve the reliability and operability of...

Read the full posting on Gimletlabs's site ↗

Research and Development

Build your edge while you search

Free tools for founders and investors, plus VC Unfiltered, our take on startups, venture and the people who build them.