Liquid AI builds efficient Liquid Foundation Models (LFMs) for on-device, edge, and cloud AI with low latency, privacy, and hardware-aware deployment.
About the role
Our Cluster Infrastructure team owns the compute environments that power foundation model training and research at Liquid AI. We are looking for a hands-on software engineer to keep our GPU clusters reliable, improve resource efficiency, and build the tooling that allows researchers to focus on model development rather than infrastructure.
What they're looking for
- Strong software engineering experience, with the ability to build production-quality infrastructure tooling and automation
- Deep knowledge of distributed systems, Linux, networking, and storage
- Experience operating a shared compute cluster or distributed training platform
- A track record of supporting production users and turning recurring failures into durable solutions
- The technical depth to partner effectively with senior research and infrastructure engineers
More about this role
Spun out of MIT CSAIL, we build general-purpose AI systems that run efficiently across deployment targets, from data center accelerators to on-device hardware, ensuring low latency, minimal memory usage, privacy, and reliability. We partner with enterprises across consumer electronics, automotive, life sciences, and financial services. We are scaling rapidly and need exceptional people to help us get there.
Our Cluster Infrastructure team owns the compute environments that power foundation model training and research at Liquid AI. We are looking for a hands-on software engineer to keep our GPU clusters reliable, improve resource efficiency, and build the tooling that allows researchers to focus on model development rather than infrastructure.
This role matters because infrastructure issues can delay training by days, while improvements in utilization, storage management, and automation can significantly increase research velocity and reduce compute costs. You will work closely with researchers and infrastructure engineers, owning problems from immediate operational response through long-term platform improvements.
Brings order to complex systems: You identify root causes and build...
Browse similar: AI jobs · AI startup jobs · Startup jobs · Remote jobs · San Francisco Bay Area