Startups

Web Scraping | Remote | Python Experience

Crowdplat · United States · On-site

← All jobs
About Crowdplat

CrowdPlat puts expert human intelligence to work at scale — platform plus services. Crowdsourcing for enterprise projects of every kind, from consulting and software builds to R&D innovation, and AI Hire for AI-powered candidate evaluation. Backed by 500 Global.

About the role

This internship offers the opportunity to build real datasets used by federal innovation and prize competitions. You will write and run Python extraction scripts against public sources including academic directories, lab pages, professional societies, conference programs, and open research APIs. The role involves delivering clean, verified CSV files with source URLs on every record, working on real datasets with real deadlines.

What they're looking for

  • Write and run extraction scripts using BeautifulSoup, Scrapy, Selenium, or Playwright
  • Use public APIs and JSON endpoints where available, in preference to page scraping
  • Handle pagination, inconsistent markup, and varied formats without losing records
  • Deliver de-duplicated CSV or Excel output with a source URL on every row
  • Validate before delivery — no malformed rows, no duplicates, no guessed email addresses
  • Stay within each site's terms of use, robots.txt, and rate limits
More about this role

**Location:** Remote (United States)

**Duration:** September–December 2026

**Hours:** 15–20 hours per week

**Overview**

This internship offers the opportunity to build real datasets used by federal innovation and prize competitions. You will write and run Python extraction scripts against public sources including academic directories, lab pages, professional societies, conference programs, and open research APIs. The role involves delivering clean, verified CSV files with source URLs on every record, working on real datasets with real deadlines.

**Key Responsibilities**

  • Write and run extraction scripts using BeautifulSoup, Scrapy, Selenium, or Playwright
  • Use public APIs and JSON endpoints where available, in preference to page scraping
  • Handle pagination, inconsistent markup, and varied formats without losing records
  • Deliver de-duplicated CSV or Excel output with a source URL on every row
  • Validate before delivery — no malformed rows, no duplicates, no guessed email addresses
  • Stay within each site's terms of use, robots.txt, and rate limits

**Required Qualifications**

  • Current student or recent graduate in CS, data science, information science, or similar
  • Working...

Read the full posting on Crowdplat's site ↗

Build your edge while you search

Free tools for founders and investors, plus VC Unfiltered, our take on startups, venture and the people who build them.