Startups · AI

Senior/Staff Software Engineer, Distributed Systems

Hedra · San Francisco · On-site

← All jobs
About Hedra

Hedra builds models that generate, perceive, and predict the visual world — and the platform that runs them in production. Backed by a16z and a16z speedrun.

About the role

We’re looking for a Senior or Staff Software Engineer with deep experience building and operating distributed production systems. You’ll work on the infrastructure underlying Hedra’s inference platform: systems that schedule and route compute, serve models efficiently, handle high-throughput workloads, and remain reliable as both traffic and the number of models we support grow.

What they're looking for

  • 5+ years of software engineering experience, with significant experience building distributed backend or infrastructure systems
  • A track record of designing and operating high-availability production services at meaningful scale
  • Strong distributed systems fundamentals, including experience reasoning about concurrency, queues, retries, failure recovery, consistency, backpressure, capacity, and system behavior under load
  • Experience debugging complex production systems across multiple layers of the stack
  • Strong judgment around architectural tradeoffs, particularly reliability, performance, complexity, and operational cost
  • Experience with CI/CD, comprehensive testing and validation strategies, monitoring, alerting, and production operations
More about this role

We build the systems that make large-scale visual inference fast, reliable, and accessible to developers. Our work spans model serving, compute infrastructure, scheduling and routing, APIs, and the developer platform that sits on top of it.

We’re a small, highly technical team in San Francisco, backed by a16z and other leading investors. Engineers at Hedra work across boundaries, own systems end to end, and have significant influence over both what we build and how we build it.

We’re looking for a Senior or Staff Software Engineer with deep experience building and operating distributed production systems.

You’ll work on the infrastructure underlying Hedra’s inference platform: systems that schedule and route compute, serve models efficiently, handle high-throughput workloads, and remain reliable as both traffic and the number of models we support grow.

The problems are often ambiguous and don’t have obvious answers. We’re looking for someone who can reason from first principles, identify bottlenecks and failure modes before they become problems, and make thoughtful tradeoffs across performance, reliability, complexity, and cost.

You do not need to come from an AI company or...

Read the full posting on Hedra's site ↗

Engineering

Build your edge while you search

Free tools for founders and investors, plus VC Unfiltered, our take on startups, venture and the people who build them.