The unified AI inference stack - from custom GPU kernels to production cloud serving on NVIDIA and AMD. 2x performance. Top open models. Open source stack. Backed by General Catalyst, Greylock and GV.
About the role
At Modular, we optimize inference from kernel to cloud on one unified stack. We are building a differentiated cloud platform that delivers state of the art inference performance from day one, then keeps getting better. As we learn the shape and patterns of each customer's workload, the platform adapts and improves performance automatically over time.
What they're looking for
- 5+ years in distributed systems or performance engineering, including experience leading or managing engineering teams
- A track record of shipping durable, reusable software tools and libraries adopted across teams and functions, and of guiding a team to do the same
- Sound judgment in evaluating technical tradeoffs and setting priorities, paired with strong communication and technical leadership skills
- The ability to translate ambiguous customer and product needs into focused engineering direction
- Creativity and curiosity in solving complex problems, a collaborative and team oriented mindset, and alignment with our culture
- Hands on background in GPU kernel programming, inference engine internals, or distributed inference architectures
More about this role
At Modular, a Qualcomm company , we’re on a mission to revolutionize AI infrastructure by systematically rebuilding the AI software stack from the ground up. Our team, made up of industry leaders and experts, is building cutting-edge, modular infrastructure that simplifies AI development and deployment. By rethinking the complexities of AI systems, we’re empowering everyone to unlock AI’s full potential and tackle some of the world’s most pressing challenges.
If you’re passionate about shaping the future of AI and creating tools that make a real difference in people’s lives, we want you on our team. You can read about our culture and careers to understand how we work and what we value.
At Modular, we optimize inference from kernel to cloud on one unified stack. We are building a differentiated cloud platform that delivers state of the art inference performance from day one, then keeps getting better. As we learn the shape and patterns of each customer's workload, the platform adapts and improves performance automatically over time.
The Performance Labs team builds the infrastructure that makes this possible at scale. We continuously apply the latest optimizations across kernels,...
Browse similar: AI jobs · AI startup jobs · Startup jobs · Remote jobs