Startups · AI

Senior AI Runtime Engineer

Modular · United States - Remote · Remote

← All jobs
About Modular

The unified AI inference stack - from custom GPU kernels to production cloud serving on NVIDIA and AMD. 2x performance. Top open models. Open source stack. Backed by General Catalyst, Greylock and GV.

About the role

ML developers today face significant friction when deploying trained models. They work in a fragmented space with incomplete, patchwork solutions that require extensive performance tuning and model-specific optimizations. At Modular, we are building the next-generation AI platform that will radically improve how developers build and deploy AI models.

What they're looking for

  • 5+ years of experience working on high-performance computing systems
  • Experience in C++ programming and complex software systems
  • Experience with CPU or GPU runtime optimizations and performance analysis on CPUs, GPUs, or AI accelerators
  • Proficiency with one or more profiling tools (CPU or GPU)
  • Creativity and curiosity for solving complex problems, a team-oriented attitude that enables you to work well with others, and alignment with our culture
  • Experience with ML graph optimizations, parallel / distributed programming, heterogeneous ML computation, and/or code generation
More about this role

At Modular, a Qualcomm company , we’re on a mission to revolutionize AI infrastructure by systematically rebuilding the AI software stack from the ground up. Our team, made up of industry leaders and experts, is building cutting-edge, modular infrastructure that simplifies AI development and deployment. By rethinking the complexities of AI systems, we’re empowering everyone to unlock AI’s full potential and tackle some of the world’s most pressing challenges.

If you’re passionate about shaping the future of AI and creating tools that make a real difference in people’s lives, we want you on our team. You can read about our culture and careers to understand how we work and what we value.

ML developers today face significant friction when deploying trained models. They work in a fragmented space with incomplete, patchwork solutions that require extensive performance tuning and model-specific optimizations. At Modular, we are building the next-generation AI platform that will radically improve how developers build and deploy AI models.

A core part of this offering is a platform that enables customers to achieve state-of-the-art performance across model families and frameworks. As an...

Read the full posting on Modular's site ↗

Build your edge while you search

Free tools for founders and investors, plus VC Unfiltered, our take on startups, venture and the people who build them.