Backed by Y Combinator.
About the role
This is a general application for candidates who are unsure which research focus - QC Automation , Benchmarks , or Synthetic Data - they would be a fit for. We would love to meet you and figure it out together. However, if you already have a focus in mind, please apply to only that application .
What they're looking for
- Proficiency in Python, Docker, and Linux environments
- Experience working on benchmarks and evals - you can reason about what makes a task realistic, a rubric reliable, an environment usable, and a trajectory useful for RL training
- Strong attention to detail and the ability to spot subtle inconsistencies in data, model behavior, or task design
- Experience building tools, pipelines, experiments, or infrastructure without a fully prescribed roadmap
- Early-stage startup experience with ability to work independently in fast-paced environments
More about this role
HUD is building infrastructure to create RL training data and evals for frontier AI agents, as well as a marketplace to sell these to frontier labs through the HUD marketplace. Our platform is used by frontier labs, Fortune 500 companies, and startups. We’ve raised $16M from top VCs and were YC W25.
This is a general application for candidates who are unsure which research focus - QC Automation , Benchmarks , or Synthetic Data - they would be a fit for. We would love to meet you and figure it out together. However, if you already have a focus in mind, please apply to only that application .
We're looking for Research Engineers to build the technical foundation for training and evaluating frontier AI agents. You’ll build the systems for creating new environments, improve data quality, and translate real-world workflows into tasks and benchmarks.
Build systems for creating, running, evaluating, and improving agent training environments
Design experiments to understand model behavior, agent failure modes, and data quality issues
Develop tools that help researchers, engineers, and data vendors create higher-quality tasks, trajectories, and feedback loops
Work across the full lifecycle...
Browse similar: AI jobs · AI startup jobs · Startup jobs · San Francisco Bay Area