Startups · AI

Hardware Failure Analysis Engineer - Memphis

xAI · Southaven, MS; Memphis, TN · On-site

← All jobs
About xAI

SpaceXAI builds Grok — frontier AI models for reasoning, voice, image generation, and more. Build with the Grok API. Backed by Lightspeed, Sequoia and a16z.

About the role

Determine true root cause of fleet hardware failures — component and system level — and drive fixes through vendors to the manufacturer. Turn "swap it again" into "vendor redesign / firmware fix / manufacturing escape found."

What they're looking for

  • Bachelor's degree in Systems Engineering, Electrical Engineering, Computer Science, or a related field (or equivalent experience)
  • 2+ years of experience in hardware reliability engineering, preferably in high-performance computing or data center environments
  • Proven expertise in firmware analysis, hardware specifications review, and release validation
  • Strong experience with RMA processes, including filing claims, vendor negotiations, and pushing for resolutions outside standard protocols
  • Demonstrated ability to diagnose and prove complex hardware failures, including grey or intermittent issues, using tools, logic analyzers, or diagnostic software
  • Familiarity with data center hardware components (e.g., servers, GPUs, networking equipment) and emerging technologies
More about this role

SpaceXAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company’s mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates.

Determine true root cause of fleet hardware failures — component and system level — and drive fixes through vendors to the manufacturer. Turn "swap it again" into "vendor redesign / firmware fix / manufacturing escape found."

  • Own named failure classes across GPU trays/baseboards, NIC/DPU, motherboard/PCIe switch, memory, power, thermal, cables/connectors, and rack-scale patterns.
  • Run recurrence analysis and fleet-wide defect clustering; detect...

Read the full posting on xAI's site ↗

Data Center

Build your edge while you search

Free tools for founders and investors, plus VC Unfiltered, our take on startups, venture and the people who build them.