# Staff/Principal DevOps Engineer, AI Inference at Lila Sciences

- Company: Lila Sciences
- What the company does: LILA has created the world's first Operating System for Science powered by Scientific Superintelligence™. Backed by General Catalyst.
- Company website: https://www.lila.ai/
- Type: Startups (AI role)
- Level: Principal and up
- Location: Cambridge, MA USA
- Work setup: On-site
- Posted: 2026-07-28
- Apply by: 2026-10-08
- Apply: https://job-boards.greenhouse.io/lilasciences/jobs/4248032009
- Page: https://www.1752.vc/careers/jobs/lila-sciences-staff-principal-devops-engineer-ai-inference/

## About the role

The Staff/Principal DevOps Engineer - AI Inference will drive the design, implementation, and optimization of infrastructure purpose-built for serving machine learning models at scale. This role bridges platform engineering, site reliability, and ML infrastructure, building the systems that power low-latency, high-throughput inference across GPU clusters and cloud accelerators.

## What they're looking for

- Expertise in DevOps, SRE, or Platform Engineering with significant experience operating GPU/accelerator infrastructure at scale
- Deep experience with Kubernetes for ML workloads: GPU scheduling, resource quotas, node affinity, and accelerator device management
- Strong proficiency deploying to AWS using infrastructure-as-code (Terraform, Helm) with hands-on experience managing GPU-based compute (EKS, EC2 P-series/Inf/Trn instances)
- Experience with model serving infrastructure: inference servers, request batching, KV-cache optimization, or LLM serving frameworks
- Strong understanding of networking for distributed inference: high-bandwidth interconnects, NCCL, VPC/PrivateLink, and load balancing at L4/L7
- Strong proficiency in Python for automation, tooling, and integration with ML frameworks

Tags: Software
