# Software Engineer, Infrastructure at Fal.ai

- Company: Fal.ai
- What the company does: Easiest & most cost-effective way to use Gen AI. fal.ai is how devs integrate dozens of generative media models. FLUX, Kling, Hailuo +1000 more. Backed by Bessemer, Kleiner Perkins and Sequoia.
- Company website: https://fal.ai/
- Type: Startups (AI role)
- Level: Mid level
- Location: San Francisco
- Work setup: On-site
- Pay: $180K to $250K base salary per year (USD)
- Posted: 2026-09-22
- Apply by: 2026-11-06
- Apply: https://jobs.ashbyhq.com/fal-ai/b8c5f81c-89c4-45c7-a1d9-3c190a523268
- Page: https://www.1752.vc/careers/jobs/fal-ai-software-engineer-infrastructure/

## About the role

You are a hands-on engineer who builds the software and processes that keep a large fleet of GPU servers healthy and productive. You write systems and tooling for managing 1000s of servers including provisioning, health monitoring, error detection, and recovery — and when something breaks that automation can’t fix, you drive resolution with partners.

## What they're looking for

- 3+ years experience managing bare-metal and cloud based server fleets at scale (100+ nodes)
- Strong software engineering skills in Python, you write production tooling, not scripts
- Deep Linux systems knowledge: boot process, kernel tuning, networking, storage, systemd, cgroups, namespaces, performance profiling
- Strong experience with configuration management and infrastructure-as-code: Ansible, Terraform, cloud-init
- Solid understanding of storage technologies: LVM, RAID, NVMe, NFS, Lustre or GPFS, and Linux I/O stack tuning
- Familiarity with hardware diagnostics and failure modes (GPUs, NVMe, NICs, memory)

Tags: Engineering
