Discover, create, and share music with the world. Use the latest technology to create AI music in seconds. Backed by a16z.
About the role
We are looking for a Senior Backend Engineer to lead the unification of large, highly rich, and heterogeneous datasets sourced from a wide range of external providers. These datasets are used to power our generative audio models.
What they're looking for
- Experience working with large, heterogeneous datasets from multiple providers or domains
- Strong background in entity resolution , deduplication, data unification, or related large-scale data integration techniques
- Proficiency in Python , with an emphasis on efficient, scalable data processing
- Experience with BigQuery, Google Dataflow/Apache Beam , or similar batch-processing frameworks
- Familiarity with data validation, normalization, reconciliation , and building consistent views across diverse data sources
- Ability to craft well-structured matching and decision strategies that balance accuracy, completeness, and computational efficiency
More about this role
We are looking for a Senior Backend Engineer to lead the unification of large, highly rich, and heterogeneous datasets sourced from a wide range of external providers. These datasets are used to power our generative audio models.
Your work will create the foundational dataset that powers our research by building robust, scalable systems for linking, deduplicating, reconciling, and enriching data at massive scale. This role centers on high-impact bulk ingestion and advanced data linkage . You will design the logic, algorithms, and strategies that transform many independent datasets into a unified, high-quality canonical asset used throughout the company.
You will collaborate closely with ML researchers and product teams, working with tools such as BigQuery, Dataflow/Beam, TFRecords , and—where beneficial—distributed systems frameworks like Ray . Familiarity with ML workflows using JAX or multihost training is a plus, as the datasets you produce will directly support that ecosystem.
- Build high-throughput bulk ingestion workflows to integrate datasets from multiple external providers.
- Design and implement scalable entity-resolution solutions, including record linking,...
Browse similar: Startup jobs