Startups · AI

Research Engineer, Benchmarks

HUD · San Francisco · On-site

← All jobs
About HUD

Backed by Y Combinator.

About the role

We’re looking for Research Engineers to build high-quality benchmarks for evaluating frontier agents on domain-specific tasks. You’ll build benchmarks that are technically rigorous, practically useful, and credible to frontier labs. Own the design, implementation, and quality of HUD’s internal agent benchmarks

What they're looking for

  • Proficiency in Python, Docker, and Linux environments
  • Published papers or written technical blogs on relevant topics such as public benchmarks and their limitations, model failure modes, etc. - please link in your application
  • Strong understanding of what a “good benchmark” means and what makes one realistic, reliable, and useful
  • Experience working on environments and evals
  • Curiosity and ability to truly understand how workflows in various domains work
More about this role

HUD is building infrastructure to create RL training data and evals for frontier AI agents, as well as a marketplace to sell these to frontier labs through the HUD marketplace. Our platform is used by frontier labs, Fortune 500 companies, and startups. We’ve raised $16M from top VCs and were YC W25.

We’re looking for Research Engineers to build high-quality benchmarks for evaluating frontier agents on domain-specific tasks. You’ll build benchmarks that are technically rigorous, practically useful, and credible to frontier labs.

Own the design, implementation, and quality of HUD’s internal agent benchmarks

Work with subject-matter experts to define tasks and create domain-specific benchmarks that evaluate agents on realistic workflows

Build infrastructure to reliably run models and agents against benchmark tasks

Develop metrics and analyses to understand benchmark difficulty, reliability, and failure modes

Validate whether benchmark performance correlates with real-world evals, customer needs, and lab expectations

Write clear documentation and benchmark reports that make results legible and credible to technical audiences

Proficiency in Python, Docker, and Linux...

Read the full posting on HUD's site ↗

Engineering & Research

Build your edge while you search

Free tools for founders and investors, plus VC Unfiltered, our take on startups, venture and the people who build them.