# Member of Technical Staff - Distributed Training Engineer at Liquid AI

- Company: Liquid AI
- What the company does: Liquid AI builds efficient Liquid Foundation Models (LFMs) for on-device, edge, and cloud AI with low latency, privacy, and hardware-aware deployment.
- Company website: https://www.liquid.ai
- Type: Startups (AI role)
- Level: Senior
- Location: San Francisco
- Work setup: Remote
- Posted: 2025-07-29
- Apply by: 2026-10-08
- Apply: https://jobs.ashbyhq.com/liquid-ai/a25b97f4-02ee-4453-a2e1-f8d5cfe2c4b4
- Page: https://www.1752.vc/careers/jobs/liquid-ai-member-of-technical-staff-distributed-training-engineer/

## About the role

Our Training Infrastructure team is building the distributed systems that power our next-generation Liquid Foundation Models. As we scale, we need to design, implement, and optimize the infrastructure that enables large-scale training. This is a high-ownership training systems role focused on runtime/performance/reliability (not a general platform/SRE role). You’ll work on a small team with fast feedback loops, building critical systems from the ground up rather than inheriting mature infrastructure.

## What they're looking for

- Hands-on experience building distributed training infrastructure (PyTorch Distributed DDP/FSDP, DeepSpeed ZeRO, Megatron-LM TP/PP)
- Experience diagnosing performance bottlenecks and failure modes (profiling, NCCL/collectives issues, hangs, OOMs, stragglers)
- Understanding of hardware accelerators and networking topologies
- Experience optimizing data pipelines for ML workloads

Tags: Research & Engineering
