Startups · AI

Inference Engineering and Product Lead

Modal · San Francisco · On-site

← All jobs
About Modal

Bring your own code, and run CPU, GPU, and data-intensive compute at scale. The serverless platform for AI and data teams. Backed by Accel, General Catalyst and AI Grant.

About the role

Modal's LLM inference platform delivers frontier performance for open-source models with best-in-class elasticity and developer experience, made in part possible by our custom runtime with GPU memory snapshots and multi-cloud substrate .

What they're looking for

  • 10+ years of industry experience, including 3+ years in a leadership role
  • Track record building high-performance systems at scale
  • Strong background in cloud infrastructure
  • Deep knowledge of low-level OS foundations (Linux kernel, file systems, containers, etc.)
  • Nice to have: Experience working with LLM inference in production and familiarity with underlying concepts like engines, kernels, routing, KV cache management and speculative decoding
More about this role

Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now.

Our customers include category-defining companies like Lovable , Ramp , Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale.

We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September.

Our team includes creators of popular open-source projects (e.g., Seaborn , Luig i ), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience.

Modal's LLM inference platform delivers frontier performance for open-source models with best-in-class elasticity and developer experience, made in part possible...

Read the full posting on Modal's site ↗

Engineering

Build your edge while you search

Free tools for founders and investors, plus VC Unfiltered, our take on startups, venture and the people who build them.