Startups · AI

Senior Inference Reliability Engineer

Parasail · San Mateo · On-site

← All jobs
About Parasail

Parasail provides AI infrastructure through a global network of on-demand GPU compute resources, enabling organizations to run and scale artificial intelligence workloads without managing physical hardware. Backed by Kindred Venture Capitals.

About the role

The Senior/Staff Inference Reliability Engineer will own the end-to-end reliability and production performance of customer inference workloads. This role sits at the intersection of inference platform engineering, LLM performance, and infrastructure reliability.

What they're looking for

  • 5+ years of experience in production engineering, site reliability engineering, infrastructure engineering, distributed systems, ML infrastructure, database reliability, or performance engineering
  • Demonstrated ownership of a critical production service or workload
  • Experience diagnosing complex latency, throughput, capacity, or reliability problems across multiple system layers
  • Strong software-engineering ability beyond infrastructure configuration and CI/CD automation
  • Hands-on experience building observability, automation, diagnostic tooling, or production safeguards
  • Strong communication and technical leadership skills, including the ability to coordinate incident resolution across engineering teams
More about this role

Parasail is redefining AI infrastructure by enabling seamless deployment across a distributed network of GPUs, optimizing for cost, performance, and flexibility. Our mission is to empower AI developers with a fast, cost-efficient, and scalable cloud experience—free from vendor lock-in and designed for the next generation of AI workloads.

The Senior/Staff Inference Reliability Engineer will own the end-to-end reliability and production performance of customer inference workloads. This role sits at the intersection of inference platform engineering, LLM performance, and infrastructure reliability.

You will ensure that customer endpoints meet expectations for availability, latency, throughput, quality, and cost. When an endpoint degrades, you will follow the problem across the entire serving path—from APIs, routing, scheduling, and autoscaling through model servers, GPUs, networking, and underlying infrastructure—and drive it through resolution.

This is not a traditional DevOps role focused only on clusters and deployments. It is a production systems role for an engineer who enjoys investigating ambiguous performance problems, building diagnostic tooling, and turning recurring...

Read the full posting on Parasail's site ↗

Software Engineering

Build your edge while you search

Free tools for founders and investors, plus VC Unfiltered, our take on startups, venture and the people who build them.