Startups · AI

Site Reliability Engineer

Blaxel · San Francisco · On-site

← All jobs
About Blaxel

Blaxel is the infrastructure foundation for autonomous agents: isolated microVMs that boot in milliseconds and resume in ~25ms, persistent shared memory, and programmable networking. Backed by Y Combinator.

About the role

We're looking for a world-class Site Reliability Engineer to ensure the reliability, performance, and scalability of our AI infrastructure platform. You’ll be building and operating the core systems that power agentic AI at scale. Your mission: keep our ultra-low-latency, stateful, serverless compute engine rock-solid as we serve billions of agent requests for the most sophisticated AI teams in the world.

What they're looking for

  • Deeply technical by default: Fluent across systems, cloud, networking, and distributed computing. You love debugging real failures, not theoretical ones
  • AI-fluent operator: You understand how AI systems behave under scale, their unique resource patterns, and the infrastructure challenges of agentic frameworks
  • Builder at heart: You want to invent new reliability systems—not just maintain existing ones. You thrive in a zero-to-one infra environment
  • High-velocity execution: You have a strong bias for action and a track record of shipping reliable systems quickly with excellent judgment
  • Automation-first mindset: You hate repeated manual work and instinctively reach for automation or AI-driven ops to scale yourself
  • Calm under pressure: When incidents hit, you operate with clarity, precision, and ownership
More about this role

We're looking for a world-class Site Reliability Engineer to ensure the reliability, performance, and scalability of our AI infrastructure platform.

You’ll be building and operating the core systems that power agentic AI at scale. Your mission: keep our ultra-low-latency, stateful, serverless compute engine rock-solid as we serve billions of agent requests for the most sophisticated AI teams in the world.

This role is highly technical and execution-heavy. You’ll own our reliability posture end-to-end—observability, performance tuning, incident ops, infrastructure health, and the automation systems that keep everything running smoothly. We want you to design new reliability systems, push the boundaries of automation, and continuously evolve the platform to meet the demands of next-generation AI workloads. If you're a builder who thrives on owning critical infrastructure at scale, this role is for you.

Collaborating closely with the founders, the infra team, and the dev team—and leveraging AI wherever it creates leverage—you will architect and operate the systems that keep Blaxel fast, resilient, and secure.

Architect, operate, and continuously improve the core infrastructure...

Read the full posting on Blaxel's site ↗

Engineering

Build your edge while you search

Free tools for founders and investors, plus VC Unfiltered, our take on startups, venture and the people who build them.