Startups

Member of Technical Staff, Data Engineering

Mithrl Inc. · San Francisco · On-site

← All jobs
About Mithrl Inc.

Mithrl is a software development company that builds custom workflows for NGS data on demand. Backed by Headline (formerly e.ventures).

About the role

Build and own an AI-powered ingestion & normalization pipeline to import data from a wide variety of sources — unprocessed Excel/CSV uploads, lab and instrument exports, as well as processed data from internal pipelines. Develop robust schema mapping, coercion, and conversion logic (think: units normalization, metadata standardization, variable-name harmonization, vendor-instrument quirks, plate-reader formats, reference-genome or annotation updates, batch-effect correction, etc.).

What they're looking for

  • Must-have
  • 5+ years of experience in data engineering / data wrangling with real-world tabular or semi-structured data
  • Strong fluency in Python, and data processing tools (Pandas, Polars, PyArrow, or similar)
  • Excellent experience dealing with messy Excel / CSV / spreadsheet-style data — inconsistent headers, multiple sheets, mixed formats, free-text fields — and normalizing it into clean structures
  • Comfort designing and maintaining robust ETL/ELT pipelines, ideally for scientific or lab-derived data
  • Ability to combine classical data engineering with LLM-powered data normalization / metadata extraction / cleaning
More about this role

We envision a world where novel drugs and therapies reach patients in months, not years, accelerating breakthroughs that save lives.

Mithrl is building the world’s first commercially available AI Co-Scientist—a discovery engine that empowers life science teams to go from messy biological data to novel insights in minutes. Scientists ask questions in natural language, and Mithrl answers with real analysis, novel targets, and patent-ready reports.

12X year-over-year revenue growth

Trusted by leading biotechs and big pharma across three continents

Driving real breakthroughs from target discovery to patient outcomes.

Build and own an AI-powered ingestion & normalization pipeline to import data from a wide variety of sources — unprocessed Excel/CSV uploads, lab and instrument exports, as well as processed data from internal pipelines.

Develop robust schema mapping, coercion, and conversion logic (think: units normalization, metadata standardization, variable-name harmonization, vendor-instrument quirks, plate-reader formats, reference-genome or annotation updates, batch-effect correction, etc.).

Use LLM-driven and classical data-engineering tools to structure “semi-structured” or messy...

Read the full posting on Mithrl Inc.'s site ↗

Engineering

Build your edge while you search

Free tools for founders and investors, plus VC Unfiltered, our take on startups, venture and the people who build them.