Parasail provides AI infrastructure through a global network of on-demand GPU compute resources, enabling organizations to run and scale artificial intelligence workloads without managing physical hardware. Backed by Kindred Venture Capitals.
About the role
At Parasail, reliability is an engineering problem that spans the entire stack. A GPU fails. A provider goes down. Traffic spikes. Customers still expect their inference to work. We’re hiring Site Reliability Engineers to build the systems that make that possible. You’ll own infrastructure across our global GPU fleet, write software that automates operations, and make the platform better at detecting, surviving, and recovering from failures.
What they're looking for
- Experience building and operating production infrastructure or distributed systems, with real ownership of reliability
- Strong Linux fundamentals and practical knowledge of networking, storage, and containers
- Hands-on experience running Kubernetes in production
- The ability to write maintainable software and automation to solve infrastructure problems
- A systematic approach to debugging problems that cross application, cluster, network, and hardware boundaries
- Good judgment about when to move quickly, when to simplify, and where reliability matters most
More about this role
AI companies need inference that’s fast, reliable, and economical at scale. Parasail delivers it. We’re building an enterprise-grade inference cloud for open-weight models where customers pay for the tokens they use, and we handle everything required to serve them.
Behind one OpenAI-compatible API, we pool GPU capacity from providers around the world and continuously optimize where and how workloads run. That means turning a changing mix of hardware, networks, and infrastructure into a service customers can trust.
We’ve raised a $32 million Series A, and we’re scaling beyond trillions of tokens a day. You’ll have the ownership and reach to shape how we get there.
At Parasail, reliability is an engineering problem that spans the entire stack. A GPU fails. A provider goes down. Traffic spikes. Customers still expect their inference to work.
We’re hiring Site Reliability Engineers to build the systems that make that possible. You’ll own infrastructure across our global GPU fleet, write software that automates operations, and make the platform better at detecting, surviving, and recovering from failures.
You’ll work directly with infrastructure, platform, and inference engineers in a...
Browse similar: Startup jobs · San Francisco Bay Area