Startups · AI

Data Center Hardware Quality & Reliability Engineer

OpenAI · San Francisco · Remote

← All jobs
About OpenAI

Backed by Greylock, Insight and Khosla.

About the role

Own the end-to-end data-center hardware quality and reliability loop for OpenAI’s 3P infrastructure and 1P current and next-gen platforms. Turn field failures into quantified risk, fast containment, verified root cause, improved MQE/NPI and manufacturing-test coverage, accurate spares forecasts, and upstream changes that prevent recurrence. The first hire must combine practical hardware/system understanding, reliability engineering, data fluency, and cross-functional technical leadership at data-center scale. Found on 1752vc Careers, the job board for startup and VC roles.

What they're looking for

  • BS in electrical, mechanical, computer, materials, reliability engineering, physics, or equivalent experience, MS preferred
  • 8+ years in hardware quality/reliability, server/rack systems, or mission-critical infrastructure, 3+ years owning field-failure, RMA, or CAPA outcomes
  • Solid working understanding of hardware and system architecture across board, tray, rack, firmware, telemetry, manufacturing test, and fleet behavior, deep expertise in every subsystem is not required
  • Reliability statistics: censored life data, Weibull/Poisson/binomial methods, confidence bounds, MTBF/MTTR, and reliability growth
  • Hands-on FMEA/FTA, accelerated or reliability-demonstration testing, 8D/CAPA, FA, and corrective-action verification
  • Working proficiency with SQL and Python/R or equivalent analytics tools
More about this role

Own the end-to-end data-center hardware quality and reliability loop for OpenAI’s 3P infrastructure and 1P current and next-gen platforms. Turn field failures into quantified risk, fast containment, verified root cause, improved MQE/NPI and manufacturing-test coverage, accurate spares forecasts, and upstream changes that prevent recurrence.

The first hire must combine practical hardware/system understanding, reliability engineering, data fluency, and cross-functional technical leadership at data-center scale.

• Build and govern the field-quality data model across telemetry, tickets, RMA/repair, FA, firmware, configuration, supplier, and manufacturing genealogy.

• Define AFR, ASR, DPPM, MTBF/MTTR, repeat-repair, NTF, repair-cycle-time, and forecast-versus-actual metrics with explicit denominators and uncertainty.

• Provide fleet-level macro views and unit/FRU/cohort-level micro views; detect shifts and bound affected populations.

• Lead systemic field-failure triage, containment, failure analysis, 8D/CAPA, risk assessment, corrective-action verification, and recurrence monitoring.

• Develop cohort, life-data, reliability-growth, and spare-demand projections by product, FRU,...

Read the full posting on OpenAI's site ↗

Scaling

Build your edge while you search

Free tools for founders and investors, plus VC Unfiltered, our take on startups, venture and the people who build them.