Startups · AI

Member of Technical Staff, AI Training Infrastructure

Fireworks · San Mateo · Remote

← All jobs
About Fireworks

Fireworks’ state of the art training and inference platform take you beyond the frontier, transforming open models into your specialized intelligence. Backed by Bessemer, Index and Lightspeed.

About the role

As a Training Infrastructure Engineer, you'll design, build, and optimize the infrastructure that powers our large-scale model training operations. Your work will be essential to developing high-performance AI training infrastructure. You'll collaborate with AI researchers and engineers to create robust training pipelines, optimize distributed training workloads, and ensure reliable model development. Design and implement scalable infrastructure for large-scale model training workloads

What they're looking for

  • Bachelor's degree in Computer Science, Computer Engineering, or related field, or equivalent practical experience
  • 3+ years of experience with distributed systems and ML infrastructure
  • Experience with PyTorch
  • Proficiency in cloud platforms (AWS, GCP, Azure)
  • Experience with containerization, orchestration (Kubernetes, Docker)
  • Knowledge of distributed training techniques (data parallelism, model parallelism, FSDP)
More about this role

Fireworks is the platform for specialized intelligence, enabling companies to build, train, and serve AI models tailored to their own data, workflows, and products. Founded by the team behind PyTorch and backed by AMD, Atreides, Benchmark Capital, Index Ventures, Lightspeed, NVIDIA, Sequoia Capital, and TCV, Fireworks powers production AI with hundreds of state-of-the-art open models across text, image, embedding, audio, and multimodal workloads. Today, Fireworks is a Series D company valued at $17.5 billion, bringing together an ambitious, collaborative team that's building the future of enterprise AI.

As a Training Infrastructure Engineer, you'll design, build, and optimize the infrastructure that powers our large-scale model training operations. Your work will be essential to developing high-performance AI training infrastructure. You'll collaborate with AI researchers and engineers to create robust training pipelines, optimize distributed training workloads, and ensure reliable model development.

Design and implement scalable infrastructure for large-scale model training workloads

Develop and maintain distributed training pipelines for LLMs and multimodal models

Optimize...

Read the full posting on Fireworks's site ↗

Engineering

Build your edge while you search

Free tools for founders and investors, plus VC Unfiltered, our take on startups, venture and the people who build them.