Make intelligence open and accessible to all. Backed by Battery, Lightspeed and Sequoia.
About the role
Reflection's Data team builds the training corpora our frontier models learn from. Before a model can learn anything, the data has to be found, fetched, extracted, and delivered reliably, responsibly, and at enormous scale. The ingestion layer is the machinery that turns the open web, licensed corpora, and other large-scale sources into well-structured, versioned, auditable datasets for pre-training.
What they're looking for
- Experience building, mentoring, and growing data or infrastructure engineering teams while staying technically hands-on. (Comfortable growing into leading a larger team quickly if you haven't managed at that scale before.)
- Deep experience building web-scale data acquisition or ingestion systems, with real ownership of production-grade pipelines at multi-TB to PB scale
- Strong coding ability and the credibility to earn the technical trust of a strong team
- Deep expertise in at least one of: web crawling & acquisition , large-scale extraction & ingestion pipelines , or data lakes / corpus storage & delivery with working knowledge across the others, and the ability to learn the rest
- Fluency with the modern large-scale data toolkit: distributed compute (Ray, Beam, Spark), orchestration (Airflow, Prefect), formats (Parquet, JSONL, WARC), and object-store / data-lake architectures
- Familiarity with how LLMs are trained and evaluated, and an intuition for what makes data useful for training, comfortable designing experiments and using proxy quality signals to guide system improvements
More about this role
Reflection is a research lab making intelligence open and accessible for everyone to use, customize, and build on. We build open models that let anyone control their intelligence and help shape the future of AI. Our mission: make intelligence open and accessible to all.
Reflection's Data team builds the training corpora our frontier models learn from. Before a model can learn anything, the data has to be found, fetched, extracted, and delivered reliably, responsibly, and at enormous scale. The ingestion layer is the machinery that turns the open web, licensed corpora, and other large-scale sources into well-structured, versioned, auditable datasets for pre-training.
As Data Ingestion Lead, you'll provide front-line leadership of the team that builds this layer, spanning all three of its pillars web crawl , data ingestion pipelines , and data lakes . You'll build, mentor, and grow a team of data ingestion engineers, guide the technical and architectural decisions across crawling, extraction, and corpus storage/delivery, and work closely with the pre-training research, data quality, and data partnerships teams that depend on what you ship. You'll stay close enough to the stack to...
Browse similar: AI jobs · AI startup jobs · Startup jobs · San Francisco Bay Area