Startups · AI

Site Reliability Engineer - Memphis

xAI · Southaven, MS; Memphis, TN · On-site

← All jobs
About xAI

SpaceXAI builds Grok — frontier AI models for reasoning, voice, image generation, and more. Build with the Grok API. Backed by Lightspeed, Sequoia and a16z.

About the role

As a Site Reliability Engineer focused on campus reliability, you will design what the campus watches and trusts, technically command cross-discipline SEVs, and build the guardrails that make the next incident smaller. You are the connective tissue across compute, network, storage, power, and cooling. This role demands calm incident leadership, fleet-scale observability judgment, and the ability to drive reliability work across software and facility boundaries.

What they're looking for

  • Bachelor's degree in Systems Engineering, Computer Science, Electrical Engineering, or a related field (or equivalent experience)
  • Proven large-scale incident command experience and calm technical leadership on a bridge
  • Demonstrated monitoring and observability design at fleet or campus scale, including alert hygiene, suppression, and signal quality
  • Experience working across at least two of: compute, network, storage, power, and cooling / facilities telemetry
  • Experience writing and operating playbooks or runbooks with a 24/7 operations, control room, or NOC partner
  • Proficiency in scripting (Python, Bash) for automation and analysis, plus general experience in at least one systems language (C, C++, Java, Go, Rust, or similar). Not required to be expert in all of them
More about this role

SpaceXAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company’s mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates.

As a Site Reliability Engineer focused on campus reliability, you will design what the campus watches and trusts, technically command cross-discipline SEVs, and build the guardrails that make the next incident smaller. You are the connective tissue across compute, network, storage, power, and cooling. This role demands calm incident leadership, fleet-scale observability judgment, and the ability to drive reliability work across software and...

Read the full posting on xAI's site ↗

Data Center

Build your edge while you search

Free tools for founders and investors, plus VC Unfiltered, our take on startups, venture and the people who build them.