# Inference Optimization Manager at Modular

- Company: Modular
- What the company does: The unified AI inference stack - from custom GPU kernels to production cloud serving on NVIDIA and AMD. 2x performance. Top open models. Open source stack. Backed by General Catalyst, Greylock and GV.
- Company website: https://www.modular.com/
- Type: Startups (AI role)
- Level: Senior
- Location: United States - Remote
- Work setup: Remote
- Posted: 2026-06-18
- Apply by: 2026-10-08
- Apply: https://jobs.gem.com/modular/am9icG9zdDof2wqh51asrICvbEH2f4vB
- Page: https://www.1752.vc/careers/jobs/modular-inference-optimization-manager/

## About the role

At Modular, we optimize inference from kernel to cloud on one unified stack. We are building a differentiated cloud platform that delivers state of the art inference performance from day one, then keeps getting better. As we learn the shape and patterns of each customer's workload, the platform adapts and improves performance automatically over time.

## What they're looking for

- 5+ years in distributed systems or performance engineering, including experience leading or managing engineering teams
- A track record of shipping durable, reusable software tools and libraries adopted across teams and functions, and of guiding a team to do the same
- Sound judgment in evaluating technical tradeoffs and setting priorities, paired with strong communication and technical leadership skills
- The ability to translate ambiguous customer and product needs into focused engineering direction
- Creativity and curiosity in solving complex problems, a collaborative and team oriented mindset, and alignment with our culture
- Hands on background in GPU kernel programming, inference engine internals, or distributed inference architectures

