Startups

Site Reliability Engineer (High Performance Computing)

SpaceX · Hawthorne, CA · On-site

← All jobs
About SpaceX

SpaceX designs, manufactures and launches advanced rockets and spacecraft. The company was founded in 2002 to revolutionize space technology, with the ultimate goal of enabling people to live on other planets. Backed by Kleiner Perkins, a16z and Founders Fund.

About the role

SpaceX HPC is a shared compute platform used across the company — vehicle and structures simulation, machine learning, AI inference, and more. We support every program at SpaceX to design and operate the worlds most advanced rockets and satellites. This role exists to put a real Site Reliability Engineer operating model on these capabilities and accelerating the world class engineering at SpaceX: toil reduction, automation, observability, and a sustainable incident process. Found on 1752vc Careers, the job board for startup and VC roles.

What they're looking for

  • Bachelor's degree in computer science, engineering, math, or a scientific discipline, OR 2+ years of professional experience operating production infrastructure in lieu of a degree
  • 2+ years of experience with Linux operating systems in production
  • 2+ years of experience operating production infrastructure (servers, services, or networks), including monitoring, debugging, and repairing what you own
  • Position is based in Hawthorne, CA and is primarily on-site
  • Must be able to participate in an on-call rotation
  • Must be willing to work extended hours and weekends as needed for incidents, cluster bring-up, and time-critical failures
More about this role

SpaceX was founded under the belief that a future where humanity is out exploring the stars is fundamentally more exciting than one where we are not. Today SpaceX is actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars.

SpaceX HPC is a shared compute platform used across the company — vehicle and structures simulation, machine learning, AI inference, and more. We support every program at SpaceX to design and operate the worlds most advanced rockets and satellites. This role exists to put a real Site Reliability Engineer operating model on these capabilities and accelerating the world class engineering at SpaceX: toil reduction, automation, observability, and a sustainable incident process.

We are looking for a Site Reliability Engineer who wants to own everything from Linux machines and our Infrastructure as Code, storage, and user facing applications – the whole ecosystem as a product, not as a ticket queue. You do not need a prior HPC title. You do need production instincts — you have operated real infrastructure, you write code to delete toil, and you care about whether users can actually get work done, not just...

Read the full posting on SpaceX's site ↗

Vehicle Engineering Software & AI

Build your edge while you search

Free tools for founders and investors, plus VC Unfiltered, our take on startups, venture and the people who build them.