SpaceXAI builds Grok — frontier AI models for reasoning, voice, image generation, and more. Build with the Grok API. Backed by Lightspeed, Sequoia and a16z.
About the role
As the Director of Site Operations, you’ll own node and rack uptime for SpaceXAI's AI supercompute cluster—the most advanced of its kind. This role is the extreme owner of cluster health and customer Service Level Agreements across 5+ sites operating 24/7. You’ll lead a 250+ person organization of site managers, shift supervisors, and technicians, plus the site reliability engineering team that monitors cluster health and drives fault mitigation at scale.
What they're looking for
- Willingness to travel frequently to data center locations to support operations across sites
- Physical capability to handle data center tasks, including lifting up to 50 lbs unassisted, standing for long periods, and occasional ladder use
- Must be willing to work extended hours and/or weekends as needed
More about this role
SpaceXAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company’s mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates.
As the Director of Site Operations, you’ll own node and rack uptime for SpaceXAI's AI supercompute cluster—the most advanced of its kind. This role is the extreme owner of cluster health and customer Service Level Agreements across 5+ sites operating 24/7. You’ll lead a 250+ person organization of site managers, shift supervisors, and technicians, plus the site reliability engineering team that monitors cluster health and drives fault mitigation...
Browse similar: AI jobs · AI startup jobs · Startup jobs