# GPU System Reliability Engineer Lead at Cowboy Space

- Company: Cowboy Space
- What the company does: Backed by a16z, Index and NEA.
- Company website: https://www.cowboyspace.com/
- Type: Startups
- Level: Senior
- Location: San Carlos, CA or Seattle, WA
- Work setup: On-site
- Pay: $150K to $225K base salary per year (USD)
- Posted: 2026-09-21
- Apply by: 2026-11-05
- Apply: https://jobs.ashbyhq.com/cowboyspace/e224aa3e-1756-456d-8f3a-dc8abe0c0b41
- Page: https://www.1752.vc/careers/jobs/cowboy-space-gpu-system-reliability-engineer-lead/

## About the role

Deploying high-performance GPU compute in Low Earth Orbit introduces a fundamentally different fault landscape than ground-based datacenter operation. This role sits at the frontier of that problem. When a fault occurs 500km above Earth, the system must detect it, classify it, contain it, and recover from it autonomously.

## What they're looking for

- Bachelors degree in Electrical Engineering or a related discipline
- 5+ years of experience in hardware validation, platform reliability engineering, or silicon validation on server-class compute systems
- Deep understanding of CPU and GPU architecture, including memory subsystems (DDR, HBM), cache hierarchies, and interconnect fabrics (PCIe, NVLink, XGMI)
- Strong knowledge of RAS concepts: error detection and correction (ECC), fault containment, error propagation, machine check architecture (MCA/MCI), and recovery mechanisms
- Hands-on experience with fault injection methodologies at hardware, firmware, and software levels
- Familiarity with system management interfaces including BMC, IPMI, Redfish, and MCTP/PLDM

Tags: Chief Engineering & Reliability
