Fireworks’ state of the art training and inference platform take you beyond the frontier, transforming open models into your specialized intelligence. Backed by Bessemer, Index and Lightspeed.
About the role
Fireworks AI is one of the industry leaders in inference and training for open models. Open models are how the rest of the world gets to build on frontier AI without handing the keys to a single vendor, and our job is to make them fast, cheap, and dependable enough that this is a real choice. That work is systems work: GPU scheduling, kernel and runtime performance, networking, storage, Linux. We serve over 40 trillion tokens a day doing it.
What they're looking for
- Systems fundamentals: 5+ years with Linux internals, system performance troubleshooting, and networking fundamentals (TCP/IP, HTTP, gRPC)
- Software engineering: 5+ years in Python, Go, C++, or Rust, writing production-grade tools and systems code
- Cloud-native operations: Operating and debugging Kubernetes, Terraform, and Docker in high-throughput production
- Distributed systems: High-throughput control planes, microservices, or multi-region setups
- Reliability fundamentals: Fault-tolerant design, SLO/SLA management, automated failover, high-availability architecture
- Influence without authority: You can get other teams to adopt a standard through credibility and useful tooling rather than mandate
More about this role
Fireworks is the platform for specialized intelligence, enabling companies to build, train, and serve AI models tailored to their own data, workflows, and products. Founded by the team behind PyTorch and backed by AMD, Atreides, Benchmark Capital, Index Ventures, Lightspeed, NVIDIA, Sequoia Capital, and TCV, Fireworks powers production AI with hundreds of state-of-the-art open models across text, image, embedding, audio, and multimodal workloads. Today, Fireworks is a Series D company valued at $17.5 billion, bringing together an ambitious, collaborative team that's building the future of enterprise AI.
Fireworks AI is one of the industry leaders in inference and training for open models. Open models are how the rest of the world gets to build on frontier AI without handing the keys to a single vendor, and our job is to make them fast, cheap, and dependable enough that this is a real choice. That work is systems work: GPU scheduling, kernel and runtime performance, networking, storage, Linux. We serve over 40 trillion tokens a day doing it.
Reliability Engineering makes sure that platform runs dependably as it grows. You will work across cloud infrastructure, AI systems, and...
Browse similar: AI jobs · AI startup jobs · Startup jobs · Remote jobs · San Francisco Bay Area