# Web Crawling - Research Engineer at Thinking Machines

- Company: Thinking Machines
- What the company does: Connectionism: Research Blog by Thinking Machines Lab. Backed by a16z, Accel and GV.
- Company website: https://thinkingmachines.ai/
- Type: Startups (AI role)
- Level: Mid level
- Location: San Francisco
- Work setup: Remote
- Pay: $350K to $475K base salary per year (USD)
- Posted: 2026-09-24
- Apply by: 2026-11-08
- Apply: https://jobs.ashbyhq.com/thinkingmachines/a5559c53-7085-45ff-8a0f-5e47b5e822bc
- Page: https://www.1752.vc/careers/jobs/thinking-machines-web-crawling-research-engineer/

## About the role

We're hiring a Software Engineer to build and own our web-crawling systems, from distributed collection at internet scale through filtering, deduplication, and deciding what data we keep.

## What they're looking for

- 8+ years designing, building, and scaling web crawlers, scrapers, or large-scale distributed data-acquisition systems
- A track record of owning crawler or data-acquisition infrastructure at internet scale
- Strong software engineering skills in a language such as Python, Go, or Rust, with real experience in distributed systems
- Working knowledge of the practical and legal considerations of large-scale web data collection (robots.txt, rate limiting, licensing)
- Experience applying machine learning to crawl selection, extraction, or data quality classification at internet scale
- Experience setting technical direction for a crawling, data acquisition, or search infrastructure team, whether or not that was your formal title

Tags: Core Engineering
