Startups · AI

Web Crawling - Research Engineer

Thinking Machines · San Francisco · Remote

← All jobs
About Thinking Machines

Connectionism: Research Blog by Thinking Machines Lab. Backed by a16z, Accel and GV.

About the role

We're hiring a Software Engineer to build and own our web-crawling systems, from distributed collection at internet scale through filtering, deduplication, and deciding what data we keep.

What they're looking for

  • 8+ years designing, building, and scaling web crawlers, scrapers, or large-scale distributed data-acquisition systems
  • A track record of owning crawler or data-acquisition infrastructure at internet scale
  • Strong software engineering skills in a language such as Python, Go, or Rust, with real experience in distributed systems
  • Working knowledge of the practical and legal considerations of large-scale web data collection (robots.txt, rate limiting, licensing)
  • Experience applying machine learning to crawl selection, extraction, or data quality classification at internet scale
  • Experience setting technical direction for a crawling, data acquisition, or search infrastructure team, whether or not that was your formal title
More about this role

The mission of Thinking Machines is to build AI that extends human will and judgment. We are training frontier models with Inkling, developing Tinker to let people make models their own, and crafting interfaces that broaden human-AI communication. We believe the future worth building is human, and we're hiring people who want to build it.

We're hiring a Software Engineer to build and own our web-crawling systems, from distributed collection at internet scale through filtering, deduplication, and deciding what data we keep.

The ideal candidate has built and scaled a web crawler or large-scale data-acquisition systems. In this role, you'll write and own production systems: the crawler itself, the infrastructure that runs it at scale, and the pipelines that turn raw crawls into usable pretraining data. You'll work closely with our pretraining and data teams to understand what's actually moving model quality, but this is fundamentally an engineering role, not a research one.

Design and scale the web crawler and ingestion infrastructure that sources Inkling's pretraining data

Build pipelines for large-scale extraction, deduplication, and data quality filtering

Build specialized...

Read the full posting on Thinking Machines's site ↗

Core Engineering

Build your edge while you search

Free tools for founders and investors, plus VC Unfiltered, our take on startups, venture and the people who build them.