# Site Reliability Engineer - Memphis at xAI

- Company: xAI
- What the company does: SpaceXAI builds Grok — frontier AI models for reasoning, voice, image generation, and more. Build with the Grok API. Backed by Lightspeed, Sequoia and a16z.
- Company website: https://x.ai/
- Type: Startups (AI role)
- Level: Mid level
- Location: Southaven, MS; Memphis, TN
- Work setup: On-site
- Posted: 2026-09-02
- Apply by: 2026-10-17
- Apply: https://job-boards.greenhouse.io/xai/jobs/5229153007
- Page: https://www.1752.vc/careers/jobs/xai-site-reliability-engineer-memphis/

## About the role

As a Site Reliability Engineer focused on campus reliability, you will design what the campus watches and trusts, technically command cross-discipline SEVs, and build the guardrails that make the next incident smaller. You are the connective tissue across compute, network, storage, power, and cooling. This role demands calm incident leadership, fleet-scale observability judgment, and the ability to drive reliability work across software and facility boundaries.

## What they're looking for

- Bachelor's degree in Systems Engineering, Computer Science, Electrical Engineering, or a related field (or equivalent experience)
- Proven large-scale incident command experience and calm technical leadership on a bridge
- Demonstrated monitoring and observability design at fleet or campus scale, including alert hygiene, suppression, and signal quality
- Experience working across at least two of: compute, network, storage, power, and cooling / facilities telemetry
- Experience writing and operating playbooks or runbooks with a 24/7 operations, control room, or NOC partner
- Proficiency in scripting (Python, Bash) for automation and analysis, plus general experience in at least one systems language (C, C++, Java, Go, Rust, or similar). Not required to be expert in all of them

Tags: Data Center
