Backed by Y Combinator.
About the role
We’re looking for a Technical Program Manager to own complex, time-sensitive data and evaluation programs for frontier AI labs, data vendors, and internal teams. You’ll turn ambiguous technical asks into executable programs, coordinate the people and dependencies required to deliver them, and ensure the resulting data is genuinely useful for evaluating and improving LLMs and agents.
What they're looking for
- Experience owning complex technical programs in AI/ML, data operations, support engineering, applied research, forward-deployed engineering, or a similarly cross-functional environment
- Strong quantitative and qualitative judgment about data—you can move between aggregate metrics and individual examples to determine whether a dataset or evaluation is useful, reliable, and representative
- Experience maintaining and improving complex processes, especially within project delivery or data contexts
- Demonstrated ability to bring structure to ambiguous problems, maintain momentum as requirements change, and make sound tradeoffs without waiting for a perfect specification
- Strong written and verbal communication skills, including the ability to turn messy technical information into clear decisions, owners, and next steps
More about this role
HUD is building infrastructure to create RL training data and evals for frontier AI agents, as well as a marketplace to sell these to frontier labs through the HUD marketplace. Our platform is used by frontier labs, Fortune 500 companies, and startups. We’ve raised $16M from top VCs and were YC W25.
We’re looking for a Technical Program Manager to own complex, time-sensitive data and evaluation programs for frontier AI labs, data vendors, and internal teams. You’ll turn ambiguous technical asks into executable programs, coordinate the people and dependencies required to deliver them, and ensure the resulting data is genuinely useful for evaluating and improving LLMs and agents. You’ll also manage vendors and build the processes, tooling, and operating rhythms that allow HUD to deliver reliably at increasing scale.
Own data and evaluation programs end-to-end, from initial scoping and requirements gathering through production, quality assurance, delivery, and retrospective
Translate ambiguous requests from AI labs and internal teams into clear specifications such as milestones, owners, dependencies, acceptance criteria, etc.
Maintain and improve data quality procedures using both...
Browse similar: AI jobs · AI startup jobs · Startup jobs · San Francisco Bay Area