Startups

Data Engineer

Wynd · Remote · Remote

← All jobs
About Wynd

Pilotez votre commerce unifié avec la solution retail ChapsRetail : OMS, encaissement, clienteling, fidélisation. Canaux et stocks synchronisés en temps réel. Backed by Techstars.

About the role

We are seeking a Data Engineer to support and improve large-scale data pipelines and infrastructure. You’ll work across data collection, processing, transformation, validation, and delivery, with a focus on scalability, reliability, and performance. This is a hands-on role where you’ll work with distributed systems, large datasets, web scraping infrastructure, and production data workloads.

What they're looking for

  • Bachelor’s degree or equivalent work experience
  • Python (advanced) — strong grasp of async programming, multiprocessing, and writing production-grade code for long-running data jobs
  • Web scraping at scale — hands-on experience with high-volume scraping (proxies, rate limiting, anti-bot evasion). Experience with platform APIs and large media/metadata datasets (video platforms, social media)
  • Distributed data pipelines — experience designing and operating pipelines across many workers/servers using task queues (Celery, Kafka, RabbitMQ, or similar)
  • Data warehousing — practical experience with columnar/analytical warehouses, Databend, ClickHouse, or BigQuery strongly preferred, comfortable with complex analytical queries, partitioning strategies, cost-aware querying on cloud warehouses
  • Docker & Kubernetes — containerizing workloads, writing Helm charts/manifests, managing deployments, autoscaling scraping/processing workloads
More about this role

We build infrastructure that delivers massive amounts of web data to the companies training the world’s most powerful AI models.

We're the team that helps to power and support Grass, a bandwidth-sharing network that lets us operate a massive distributed crawler, giving us unique access to high-quality public web data at global scale. On top of that, we’ve built pipelines for ingesting, segmenting, and annotating billions of videos, transcripts, and audio files, powering dataset creation for frontier labs.

We’re lean, technical, and move fast. No red tape, no slow decision-making; just a team of builders pushing to expand what’s possible for open web data and AI.

We are seeking a Data Engineer to support and improve large-scale data pipelines and infrastructure. You’ll work across data collection, processing, transformation, validation, and delivery, with a focus on scalability, reliability, and performance.

This is a hands-on role where you’ll work with distributed systems, large datasets, web scraping infrastructure, and production data workloads.

Please note: This role requires a work schedule that overlaps sufficiently with EST business hours to collaborate effectively with the...

Read the full posting on Wynd's site ↗

Analytics

Build your edge while you search

Free tools for founders and investors, plus VC Unfiltered, our take on startups, venture and the people who build them.