Build what's next on the AI Native Cloud. Full-stack AI platform for inference, fine-tuning, and GPU clusters — powered by cutting-edge research. Backed by General Catalyst, Kleiner Perkins and NEA.
About the role
We’re looking for a detail-oriented Data Center Operations professional to manage and track all break/fix activities across multiple data center locations. This role acts as the central point of coordination for hardware incidents, vendor dispatches, ticket management, asset tracking, and operational reporting to ensure maximum uptime and fast issue resolution.
What they're looking for
- Experience working in data center operations, IT infrastructure, or hardware support
- Strong understanding of server, storage, and networking hardware
- Experience with ticketing systems such as ServiceNow, Jira, or Remedy
- Ability to manage multiple priorities across several sites simultaneously
- Excellent communication and organizational skills
- Familiarity with SLA management and incident escalation processes
More about this role
We’re looking for a detail-oriented Data Center Operations professional to manage and track all break/fix activities across multiple data center locations. This role acts as the central point of coordination for hardware incidents, vendor dispatches, ticket management, asset tracking, and operational reporting to ensure maximum uptime and fast issue resolution.
- Track and manage all break/fix incidents across multiple data centers
- Monitor ticket queues and ensure SLA compliance for incident response and resolution
- Coordinate with on-site technicians, remote hands teams, vendors, and engineering groups
- Maintain accurate records of failed hardware, replacements, RMAs, and repair status
- Escalate critical outages and recurring infrastructure issues to leadership and engineering teams
- Schedule and oversee maintenance windows and emergency repair activities
- Provide daily/weekly operational status reports and incident summaries
- Ensure all work follows data center operational procedures and change management policies
- Identify trends in hardware failures and recommend process improvements
- Experience working in data center operations, IT infrastructure, or hardware...
Browse similar: AI jobs · AI startup jobs · Startup jobs · San Francisco Bay Area