# Inference Optimization Engineer at Modular

- Company: Modular
- What the company does: The unified AI inference stack - from custom GPU kernels to production cloud serving on NVIDIA and AMD. 2x performance. Top open models. Open source stack. Backed by General Catalyst, Greylock and GV.
- Company website: https://www.modular.com/
- Type: Startups (AI role)
- Level: Mid level
- Location: United States - Remote
- Work setup: Remote
- Posted: 2026-06-17
- Apply by: 2026-10-08
- Apply: https://jobs.gem.com/modular/am9icG9zdDrb7o68jwVglvW4vP1xIe82
- Page: https://www.1752.vc/careers/jobs/modular-inference-optimization-engineer/

## About the role

At Modular, we optimize inference from kernel to cloud on one unified stack. We are building a differentiated cloud platform that delivers state of the art inference performance from day one, then keeps getting better. As we learn the shape and patterns of each customer's workload, the platform adapts and improves performance automatically over time.

## What they're looking for

- 5+ years of experience in distributed systems or performance engineering
- A track record of building durable, reusable software tools and libraries that are adopted across teams and functions
- Sound judgment in evaluating technical tradeoffs and setting priorities, paired with strong communication and technical leadership skills
- Creativity and curiosity in solving complex problems, a collaborative and team oriented mindset, and alignment with our culture
- Experience with GPU kernel programming, inference engine internals, or distributed inference architectures
- Experience with Kubernetes and cloud native ecosystems

