Startups · AI

Data Center Operations Coordinator

Together AI · San Francisco · On-site

← All jobs
About Together AI

Build what's next on the AI Native Cloud. Full-stack AI platform for inference, fine-tuning, and GPU clusters — powered by cutting-edge research. Backed by General Catalyst, Kleiner Perkins and NEA.

About the role

We’re looking for a detail-oriented Data Center Operations professional to manage and track all break/fix activities across multiple data center locations. This role acts as the central point of coordination for hardware incidents, vendor dispatches, ticket management, asset tracking, and operational reporting to ensure maximum uptime and fast issue resolution.

What they're looking for

  • Experience working in data center operations, IT infrastructure, or hardware support
  • Strong understanding of server, storage, and networking hardware
  • Experience with ticketing systems such as ServiceNow, Jira, or Remedy
  • Ability to manage multiple priorities across several sites simultaneously
  • Excellent communication and organizational skills
  • Familiarity with SLA management and incident escalation processes
More about this role

We’re looking for a detail-oriented Data Center Operations professional to manage and track all break/fix activities across multiple data center locations. This role acts as the central point of coordination for hardware incidents, vendor dispatches, ticket management, asset tracking, and operational reporting to ensure maximum uptime and fast issue resolution.

  • Track and manage all break/fix incidents across multiple data centers
  • Monitor ticket queues and ensure SLA compliance for incident response and resolution
  • Coordinate with on-site technicians, remote hands teams, vendors, and engineering groups
  • Maintain accurate records of failed hardware, replacements, RMAs, and repair status
  • Escalate critical outages and recurring infrastructure issues to leadership and engineering teams
  • Schedule and oversee maintenance windows and emergency repair activities
  • Provide daily/weekly operational status reports and incident summaries
  • Ensure all work follows data center operational procedures and change management policies
  • Identify trends in hardware failures and recommend process improvements
  • Experience working in data center operations, IT infrastructure, or hardware...

Read the full posting on Together AI's site ↗

Engineering

Build your edge while you search

Free tools for founders and investors, plus VC Unfiltered, our take on startups, venture and the people who build them.