Backed by 500 Global.
About the role
You’ll work on designing, building, and scaling infrastructure for serving top open-source AI models in production. This role is ideal for engineers who are already comfortable owning problems end-to-end and want to deepen their experience working on high-impact AI systems. If you’re excited about AI/ML, have built and shipped projects, and are looking to work on real systems at scale — we’d love to meet you.
What they're looking for
- Bachelor’s or Master’s degree in Computer Science, Computer Engineering, or a related field
- 1–4 years of relevant experience , including early full-time roles or research
- Strong fundamentals in data structures, algorithms, and software design
- Proficiency in Python and experience working with AI/ML frameworks (e.g., PyTorch, TensorFlow)
- Hands-on experience building, shipping, and maintaining software systems
- Familiarity with AI models, Transformers, and Diffusers
More about this role
DeepInfra is building the infrastructure layer for the next generation of AI. We believe open-source models are the future, and companies should have full control over their AI stack without being locked into proprietary providers.
Our inference platform serves trillions of tokens every week across hundreds of production workloads. We build everything from GPU infrastructure to the API layer because every millisecond matters.
We are looking for strong Software Engineers to join our team.
You’ll work on designing, building, and scaling infrastructure for serving top open-source AI models in production. This role is ideal for engineers who are already comfortable owning problems end-to-end and want to deepen their experience working on high-impact AI systems.
If you’re excited about AI/ML, have built and shipped projects, and are looking to work on real systems at scale — we’d love to meet you.
- Design, develop, and test inference solutions for state-of-the-art AI models
- Implement, optimize, and evaluate AI models using Python, C++, CUDA, and NCCL
- Own and operate production model-serving systems, including monitoring and debugging
- Build new features, improve system...
Browse similar: Startup jobs · San Francisco Bay Area