Startups · AI

Site Reliability Engineer

Beam · New York, NY, US / San Francisco, CA, US / Remote (US) · Remote

← All jobs
About Beam

AI-Native Cloud Platform. Backed by Index.

About the role

Beam is an ultrafast AI inference platform. We built a serverless runtime that launches GPU-backed containers in less than 1 second and quickly scales out to thousands of GPUs. Developers use our platform to serve apps to millions of users around the globe. We're backed by Y Combinator, Tiger Global, and prominent developer-tool founders, including the founder of Snyk and former CTO of GitHub.

What they're looking for

  • You have an instinct for hardware. You're comfortable reasoning about failure modes at the firmware and silicon level, not just the software stack above it
  • You're fluent with AI tooling. You aren’t afraid to max-out your token usage for the right spec
  • You’re comfortable debugging production issues, from triage to post-mortem
  • Enthusiasm for developer tools, cloud native technologies, and open source software
More about this role

Beam is an ultrafast AI inference platform. We built a serverless runtime that launches GPU-backed containers in less than 1 second and quickly scales out to thousands of GPUs. Developers use our platform to serve apps to millions of users around the globe. We're backed by Y Combinator, Tiger Global, and prominent developer-tool founders, including the founder of Snyk and former CTO of GitHub.

  • Own compute fleet health end to end. Build the metrics pipelines, alerting, and unified health view that tell you the true state of every GPU in production.
  • Turn deployment debugging into an automated pipeline, not a runbook. Build and own the automation that takes a compute failure from detection through triage.
  • Design the GPU qualification platform. Burn-in, performance baselining, and NPI execution for every new GPU we onboard to our platform. You define what "good" looks like before hardware goes into production.
  • Own firmware-level telemetry, log collection at scale, and the low-level access layer that repair automation and health tooling depend on.
  • You have an instinct for hardware. You're comfortable reasoning about failure modes at the firmware and silicon level, not just...

Read the full posting on Beam's site ↗

Engineering

Build your edge while you search

Free tools for founders and investors, plus VC Unfiltered, our take on startups, venture and the people who build them.