Build what's next on the AI Native Cloud. Full-stack AI platform for inference, fine-tuning, and GPU clusters — powered by cutting-edge research. Backed by General Catalyst, Kleiner Perkins and NEA.
About the role
Together AI is growing its compute footprint, and making sure new capacity meets our technical standards is an important priority for the company. Every new cluster has to clear a technical bar before it carries customer workloads, and this role owns that bar.
What they're looking for
- 5+ years in technical program or project management, infrastructure program management, or a comparable technical operations role, ideally involving hardware, data center, or large-scale compute environments
- Proven ability to run multiple complex, cross-functional workstreams to deadline, with strong organization and stakeholder management
- Hands-on comfort with data: able to write scripts or queries (for example, Python or SQL) to compare, validate, and analyze provider specifications and test results independently
- Excellent written and verbal communication, able to turn dense technical detail into clear recommendations for both engineers and executives
- Willingness to travel to provider and data center sites as needed
- This is not an engineering manager role
More about this role
Together AI is growing its compute footprint, and making sure new capacity meets our technical standards is an important priority for the company. Every new cluster has to clear a technical bar before it carries customer workloads, and this role owns that bar. As Technical Compute Qualification Manager, you will run the process that screens and qualifies prospective compute providers, taking each prospective deployment through a structured evaluation across compute, networking, storage, power, cooling, and operations.
You will coordinate various engineering partners through validation, review provider specifications and test results, and produce clear go/no-go recommendations on whether new capacity meets our standards. It is a high-impact, process-driven role for someone technical enough to know when a spec sheet does not add up, and additional diligence needs to be completed, and organized enough to drive many evaluations to closure in parallel. "You will deep-dive into critical hardware performance metrics, proactively identifying potential bottlenecks in cluster architecture before they impact our end customers training or inference workloads." Conduct diligence and work with...
Browse similar: AI jobs · AI startup jobs · Startup jobs · San Francisco Bay Area