Backed by Y Combinator.
About the role
We’re looking for a Research Engineer to build the privacy and anonymization systems that make sensitive, real-world data safe and useful for AI training. You’ll develop methods to detect and remove PII, secrets, and other sensitive information from raw data before it enters our processing and synthetic data pipelines. You’ll own the full pipeline for protecting privacy without destroying the structure and signal that make data valuable for training agents.
What they're looking for
- Strong proficiency in Python and experience building reliable production data or ML systems
- Experience with information extraction, named-entity recognition, classification, or related methods for detecting rare or sensitive content
- Strong experimental instincts and the ability to compare approaches across recall, precision, latency, cost, and downstream data utility
- An understanding of the difference between redaction, masking, pseudonymization, anonymization, and synthetic data—and when each is appropriate
- High attention to detail and the ability to reason about subtle leakage paths, edge cases, and adversarial failure modes
- Built data processing pipelines end-to-end without a fully prescribed roadmap
More about this role
HUD is building infrastructure to create RL training data and evals for frontier AI agents, as well as a marketplace to sell these to frontier labs through the HUD marketplace. Our platform is used by frontier labs, Fortune 500 companies, and startups. We’ve raised $16M from top VCs and were YC W25.
We’re looking for a Research Engineer to build the privacy and anonymization systems that make sensitive, real-world data safe and useful for AI training. You’ll develop methods to detect and remove PII, secrets, and other sensitive information from raw data before it enters our processing and synthetic data pipelines. You’ll own the full pipeline for protecting privacy without destroying the structure and signal that make data valuable for training agents.
Build systems to detect PII, quasi-identifiers, credentials, and other sensitive information and design transformations based on the data type and downstream use case
Develop and benchmark detection approaches that combine rules, statistical models, classifiers, and LLM-based methods
Build production pipelines that anonymize raw data before it enters downstream processing, training, evaluation, or synthetic data generation...
Browse similar: AI jobs · AI startup jobs · Startup jobs · San Francisco Bay Area