# Research Engineer, Benchmarks at HUD

- Company: HUD
- What the company does: Backed by Y Combinator.
- Company website: https://www.hud.ai
- Type: Startups (AI role)
- Level: Mid level
- Location: San Francisco
- Work setup: On-site
- Posted: 2026-07-13
- Apply by: 2026-10-08
- Apply: https://jobs.ashbyhq.com/hud/c5252f0e-fd1d-41cf-b803-b0ce1fe4cea1
- Page: https://www.1752.vc/careers/jobs/hud-research-engineer-benchmarks/

## About the role

We’re looking for Research Engineers to build high-quality benchmarks for evaluating frontier agents on domain-specific tasks. You’ll build benchmarks that are technically rigorous, practically useful, and credible to frontier labs. Own the design, implementation, and quality of HUD’s internal agent benchmarks

## What they're looking for

- Proficiency in Python, Docker, and Linux environments
- Published papers or written technical blogs on relevant topics such as public benchmarks and their limitations, model failure modes, etc. - please link in your application
- Strong understanding of what a “good benchmark” means and what makes one realistic, reliable, and useful
- Experience working on environments and evals
- Curiosity and ability to truly understand how workflows in various domains work

Tags: Engineering & Research
