Startups · AI

Infrastructure Operations Engineer

Lightning AI · London, England, United Kingdom; New York, New York, United States; Remote; San Francisco, California, United States;... · Remote

← All jobs
About Lightning AI

From the PyTorch Lightning creators. Own your AI, don. Backed by Index.

About the role

In this role, you’ll work hands-on with large-scale GPU environments, Linux systems, bare metal infrastructure, provisioning workflows, and platform reliability. You’ll partner closely with Infrastructure Engineering, Network Operations, and Software Platform teams to troubleshoot issues, improve operational efficiency, and build automation that reduces manual toil over time.

What they're looking for

  • 8+ years working with Linux as a server / hosting platform, extra points for Ubuntu experience
  • 5+ years experience with AWS
  • 2+ years experience with Kubernetes and strong container fundamentals
  • 2+ years experience with Terraform and Ansible
  • 2+ years with network attached storage management (via NFS, ceph, or other protocols). Extra points for experience with VAST storage systems
  • Experience with monitoring systems (Prometheus, ELK stack)
More about this role

Lightning AI is the company behind PyTorch Lightning. Founded in 2019, we build an end-to-end platform for developing, training, and deploying AI systems—designed to take ideas from research to production with less friction.

Through our merger with Voltage Park, a neocloud and AI Factory, Lightning AI combines developer-first software with cost-efficient, large-scale compute. Teams get the tools they need for experimentation, training, and production inference, with security, observability, and control built in.

We serve solo researchers, startups, and large enterprises. Lightning AI operates globally with offices in New York City, San Francisco, Seattle, and London, and is backed by Coatue, Index Ventures, Bain Capital Ventures, and Firstminute.

The people who thrive here are builders who move fast, communicate openly, take ownership, and continuously improve themselves, their teams, and our company. Here's what that looks like in practice:

  • Move with Urgency: We move quickly, make thoughtful decisions, and keep momentum. We value action over perfection and learn by shipping.
  • Take Ownership: We own outcomes, not just our individual work. We make decisions that move the company...

Read the full posting on Lightning AI's site ↗

Infrastructure Engineering

Build your edge while you search

Free tools for founders and investors, plus VC Unfiltered, our take on startups, venture and the people who build them.