# Infrastructure Operations Engineer at Lightning AI

- Company: Lightning AI
- What the company does: From the PyTorch Lightning creators. Own your AI, don. Backed by Index.
- Company website: https://lightning.ai/
- Type: Startups (AI role)
- Level: Mid level
- Location: London, England, United Kingdom; New York, New York, United States; Remote; San Francisco, California, United States;...
- Work setup: Remote
- Posted: 2026-05-20
- Apply by: 2026-10-08
- Apply: https://job-boards.greenhouse.io/lightningai/jobs/7737825003
- Page: https://www.1752.vc/careers/jobs/lightning-ai-infrastructure-operations-engineer/

## About the role

In this role, you’ll work hands-on with large-scale GPU environments, Linux systems, bare metal infrastructure, provisioning workflows, and platform reliability. You’ll partner closely with Infrastructure Engineering, Network Operations, and Software Platform teams to troubleshoot issues, improve operational efficiency, and build automation that reduces manual toil over time.

## What they're looking for

- 8+ years working with Linux as a server / hosting platform, extra points for Ubuntu experience
- 5+ years experience with AWS
- 2+ years experience with Kubernetes and strong container fundamentals
- 2+ years experience with Terraform and Ansible
- 2+ years with network attached storage management (via NFS, ceph, or other protocols). Extra points for experience with VAST storage systems
- Experience with monitoring systems (Prometheus, ELK stack)

Tags: Infrastructure Engineering
