Goodfire is an AI interpretability research lab focused on understanding and intentionally designing advanced AI systems. Backed by Lightspeed, AI Grant and Menlo.
About the role
We’re hiring a Compute Lead to own one of the most critical resources at the company: our compute for model training, research, and customers. You’ll run everything from how we allocate our compute to how we source, test, and secure capacity. You’ll make the day-to-day calls that determine how far our world-class research and product ambitions can go.
What they're looking for
- Manage Goodfire’s compute allocations across training, research, and customer workloads
- Source and secure capacity: gather availability and pricing across providers, move quickly on the best options, and take deals from first call through contract and delivery
- Own cluster acceptance testing: define acceptance criteria with our infrastructure team and run burn-in and performance testing before we sign off
- Monitor cluster health once live: track node failures and downtime, and drive fixes with providers
- Track utilization and quota usage by project, flag idle or underused capacity, and reallocate quickly so GPUs are going to our highest-priority work
- Partner closely with research and product teams to translate roadmaps into concrete compute requirements and ensure infrastructure never becomes a bottleneck
More about this role
Goodfire is an AI interpretability research lab. We build interpretability agents and infrastructure to understand, monitor, and align AI models.
We believe that AI is the most consequential technology of our time. Yet AI models are trained by trial and error today, with limited understanding of what drives their behavior. Our goal is to advance the research and technology needed to fully understand and align AI so humanity can trust superintelligence.
Goodfire is a public benefit corporation founded by researchers who helped pioneer the field of interpretability at OpenAI and Google DeepMind. We are headquartered in San Francisco and have raised over $200M from leading investors including Lightspeed, Menlo, and B Capital.
We’re hiring a Compute Lead to own one of the most critical resources at the company: our compute for model training, research, and customers. You’ll run everything from how we allocate our compute to how we source, test, and secure capacity. You’ll make the day-to-day calls that determine how far our world-class research and product ambitions can go.
We’re looking for someone who can navigate GPU scarcity and partner negotiations with rigor, and moves with...
Browse similar: AI jobs · AI startup jobs · Startup jobs · San Francisco Bay Area