# Staff Engineer, Distributed Storage and HPC & AI Infrastructure at Together AI

- Company: Together AI
- What the company does: Build what's next on the AI Native Cloud. Full-stack AI platform for inference, fine-tuning, and GPU clusters — powered by cutting-edge research. Backed by General Catalyst, Kleiner Perkins and NEA.
- Company website: https://www.together.ai/
- Type: Startups (AI role)
- Level: Senior
- Location: San Francisco
- Work setup: On-site
- Pay: $250K to $300K base salary per year (USD)
- Posted: 2026-06-04
- Apply by: 2026-10-08
- Apply: https://job-boards.greenhouse.io/togetherai/jobs/5155722007
- Page: https://www.1752.vc/careers/jobs/together-ai-staff-engineer-distributed-storage-and-hpc-and-ai-infrastructure/

## About the role

In this role, you will operate, scale, and optimize multi-petabyte storage systems purpose-built for the world’s largest AI training and inference workloads. You’ll manage and scale high-performance parallel filesystems and object stores, evaluate and integrate cutting-edge technologies such as Vast, Weka, Ceph, and Lustre, and solve the complex engineering challenges of operating at extreme throughput, low-latency data paths, and massive cluster-scale storage operations.

## What they're looking for

- 8+ years in storage engineering, managing distributed storage at multi-petabyte scale
- Proven track record deploying and operating high-performance storage for GPU/HPC clusters
- Deep Kubernetes and cloud-native storage experience in production environments
- Strong coding skills in Go and Python with demonstrated ability to build production-grade systems and tooling
- BS/MS in Computer Science, Engineering, or equivalent practical experience
- History of technical leadership: designing systems that significantly improved performance, reliability (99.999%+ uptime), or cost efficiency

Tags: Engineering
