Startups · AI

Research Engineer, Privacy and Anonymization

HUD · San Francisco · On-site

← All jobs
About HUD

Backed by Y Combinator.

About the role

We’re looking for a Research Engineer to build the privacy and anonymization systems that make sensitive, real-world data safe and useful for AI training. You’ll develop methods to detect and remove PII, secrets, and other sensitive information from raw data before it enters our processing and synthetic data pipelines. You’ll own the full pipeline for protecting privacy without destroying the structure and signal that make data valuable for training agents.

What they're looking for

  • Strong proficiency in Python and experience building reliable production data or ML systems
  • Experience with information extraction, named-entity recognition, classification, or related methods for detecting rare or sensitive content
  • Strong experimental instincts and the ability to compare approaches across recall, precision, latency, cost, and downstream data utility
  • An understanding of the difference between redaction, masking, pseudonymization, anonymization, and synthetic data—and when each is appropriate
  • High attention to detail and the ability to reason about subtle leakage paths, edge cases, and adversarial failure modes
  • Built data processing pipelines end-to-end without a fully prescribed roadmap
More about this role

HUD is building infrastructure to create RL training data and evals for frontier AI agents, as well as a marketplace to sell these to frontier labs through the HUD marketplace. Our platform is used by frontier labs, Fortune 500 companies, and startups. We’ve raised $16M from top VCs and were YC W25.

We’re looking for a Research Engineer to build the privacy and anonymization systems that make sensitive, real-world data safe and useful for AI training. You’ll develop methods to detect and remove PII, secrets, and other sensitive information from raw data before it enters our processing and synthetic data pipelines. You’ll own the full pipeline for protecting privacy without destroying the structure and signal that make data valuable for training agents.

Build systems to detect PII, quasi-identifiers, credentials, and other sensitive information and design transformations based on the data type and downstream use case

Develop and benchmark detection approaches that combine rules, statistical models, classifiers, and LLM-based methods

Build production pipelines that anonymize raw data before it enters downstream processing, training, evaluation, or synthetic data generation...

Read the full posting on HUD's site ↗

Engineering & Research

Build your edge while you search

Free tools for founders and investors, plus VC Unfiltered, our take on startups, venture and the people who build them.