Startups

Senior Site Reliability Engineer

Parasail · San Mateo · On-site

← All jobs
About Parasail

Parasail provides AI infrastructure through a global network of on-demand GPU compute resources, enabling organizations to run and scale artificial intelligence workloads without managing physical hardware. Backed by Kindred Venture Capitals.

About the role

At Parasail, reliability is an engineering problem that spans the entire stack. A GPU fails. A provider goes down. Traffic spikes. Customers still expect their inference to work. We’re hiring Site Reliability Engineers to build the systems that make that possible. You’ll own infrastructure across our global GPU fleet, write software that automates operations, and make the platform better at detecting, surviving, and recovering from failures.

What they're looking for

  • Experience building and operating production infrastructure or distributed systems, with real ownership of reliability
  • Strong Linux fundamentals and practical knowledge of networking, storage, and containers
  • Hands-on experience running Kubernetes in production
  • The ability to write maintainable software and automation to solve infrastructure problems
  • A systematic approach to debugging problems that cross application, cluster, network, and hardware boundaries
  • Good judgment about when to move quickly, when to simplify, and where reliability matters most
More about this role

AI companies need inference that’s fast, reliable, and economical at scale. Parasail delivers it. We’re building an enterprise-grade inference cloud for open-weight models where customers pay for the tokens they use, and we handle everything required to serve them.

Behind one OpenAI-compatible API, we pool GPU capacity from providers around the world and continuously optimize where and how workloads run. That means turning a changing mix of hardware, networks, and infrastructure into a service customers can trust.

We’ve raised a $32 million Series A, and we’re scaling beyond trillions of tokens a day. You’ll have the ownership and reach to shape how we get there.

At Parasail, reliability is an engineering problem that spans the entire stack. A GPU fails. A provider goes down. Traffic spikes. Customers still expect their inference to work.

We’re hiring Site Reliability Engineers to build the systems that make that possible. You’ll own infrastructure across our global GPU fleet, write software that automates operations, and make the platform better at detecting, surviving, and recovering from failures.

You’ll work directly with infrastructure, platform, and inference engineers in a...

Read the full posting on Parasail's site ↗

Software Engineering

Build your edge while you search

Free tools for founders and investors, plus VC Unfiltered, our take on startups, venture and the people who build them.