# Software Engineer, Site Reliability at Fal.ai

- Company: Fal.ai
- What the company does: Easiest & most cost-effective way to use Gen AI. fal.ai is how devs integrate dozens of generative media models. FLUX, Kling, Hailuo +1000 more. Backed by Bessemer, Kleiner Perkins and Sequoia.
- Company website: https://fal.ai/
- Type: Startups (AI role)
- Level: Mid level
- Location: Remote - Global
- Work setup: Remote
- Posted: 2026-03-13
- Apply by: 2026-10-08
- Apply: https://jobs.ashbyhq.com/fal-ai/7fbe1d25-29fb-41bb-bd52-4c8f7f277201
- Page: https://www.1752.vc/careers/jobs/fal-ai-software-engineer-site-reliability-2/

## About the role

Own and operate our Kubernetes infrastructure: cluster lifecycle, upgrades, networking, and multi-tenant isolation for customer workloads Build and maintain CI/CD pipelines and deployment infrastructure

## What they're looking for

- 5+ years experience in managing critical production systems and software development workflows
- Strong production experience setting up and operating Kubernetes at scale, using infrastructure-as-code (Terraform, Ansible)
- Deep knowledge of Linux networking, container networking (CNI plugins, VXLAN, BGP), and DNS
- Experience building CI/CD systems and GitOps workflows (FluxCD, ArgoCD)
- Proficiency in Python and either Go or Bash for tooling and automation
- Strong experience with logging, monitoring and alerting (Prometheus, Grafana, Loki, Thanos, VictoriaMetrics, Datadog)

Tags: Engineering
