Make intelligence open and accessible to all. Backed by Battery, Lightspeed and Sequoia.
About the role
Design, build, and operate large-scale GPU infrastructure for high-throughput model inference and mid-training workloads. Develop systems that power synthetic data generation and reinforcement learning pipelines at scale.
What they're looking for
- Experience deploying and operating large-scale GPU systems for inference or model serving
- Several years of hands-on experience building and running production infrastructure
- Strong understanding of GPU performance characteristics and optimization techniques
- Experience working with modern inference frameworks such as SGLang, Megatron, or similar high-performance LLM runtimes
- Familiarity with distributed reinforcement learning infrastructure or rollout generation systems
- Experience optimizing throughput for large-scale model execution workloads
More about this role
Reflection is a research lab making intelligence open and accessible for everyone to use, customize, and build on. We build open models that let anyone control their intelligence and help shape the future of AI. Our mission: make intelligence open and accessible to all.
Design, build, and operate large-scale GPU infrastructure for high-throughput model inference and mid-training workloads.
Develop systems that power synthetic data generation and reinforcement learning pipelines at scale.
Build high-performance inference platforms capable of serving and evaluating models across thousands of GPUs.
Optimize throughput, latency, and GPU utilization for large language model inference and rollout workloads.
Build infrastructure that supports reinforcement learning pipelines, including large-scale rollout generation, evaluation, and policy improvement loops.
Work closely with research teams to support distributed RL workloads and large-scale model evaluation infrastructure.
Improve performance of model execution through kernel-level optimization, model parallelism strategies, and GPU runtime improvements.
Develop distributed systems that enable large-scale synthetic data generation and...
Browse similar: AI jobs · AI startup jobs · Startup jobs · San Francisco Bay Area