# Software Engineer, Site Reliability at Fal.ai

- Company: Fal.ai
- What the company does: Easiest & most cost-effective way to use Gen AI. fal.ai is how devs integrate dozens of generative media models. FLUX, Kling, Hailuo +1000 more. Backed by Bessemer, Kleiner Perkins and Sequoia.
- Company website: https://fal.ai/
- Type: Startups (AI role)
- Level: Mid level
- Location: San Francisco
- Work setup: On-site
- Pay: $180K to $250K base salary per year (USD)
- Posted: 2026-02-23
- Apply by: 2026-10-08
- Apply: https://jobs.ashbyhq.com/fal-ai/00e59876-de9f-411d-a7e9-3a001d32e691
- Page: https://www.1752.vc/careers/jobs/fal-ai-software-engineer-site-reliability/

## About the role

You are a seasoned SRE who keeps production infrastructure running at scale. You own the reliability and availability of customer-facing systems — from Kubernetes clusters to deployment pipelines to the networking layer that connects it all. You think in SLOs, automate ruthlessly, and treat every incident as a chance to make the system better. Own and operate our Kubernetes infrastructure: cluster lifecycle, upgrades, networking, and multi-tenant isolation for customer workloads

## What they're looking for

- 5+ years experience in managing critical production systems and software development workflows
- Strong production experience setting up and operating Kubernetes at scale, using infrastructure-as-code (Terraform, Ansible)
- Deep knowledge of Linux networking, container networking (CNI plugins, VXLAN, BGP), and DNS
- Experience building CI/CD systems and GitOps workflows (FluxCD, ArgoCD)
- Proficiency in Python and either Go or Bash for tooling and automation
- Strong experience with logging, monitoring and alerting (Prometheus, Grafana, Loki, Thanos, VictoriaMetrics, Datadog)

Tags: Engineering
