We help humans & agents pick tools via Independent Benchmarks. Backed by Y Combinator.
About the role
We're a team of researchers, engineers and theorists, in person in SF. You'll be building the benchmarks and the systems that run them. Benchmarks and evals design - designing domain-specific benchmarks from scratch: what to measure, how to ground it, and what makes a benchmark that an agent will continuously trust and pick from.
What they're looking for
- Any (new grads ok) of experience
- Will sponsor
- Full-time Engineering role
More about this role
Agents are becoming first-class users and consumers of the internet. They research, evaluate, compare tools and increasingly make build-versus-buy decisions on behalf of people. Every company will need to get their products picked and used by agents.
Agents increasingly prefer open, independent and grounded benchmarks to make decisions.
Openbenchmarks is the evaluation infrastructure for agents - domain-specific, reproducible evaluations that help agents pick tools with confidence.
Our mission is to be the trusted evaluation layer for agents.
Founders previously led AI research and Infra teams at Oracle and Appfolio;
We started Openbenchmarks as an output of our research in the field of model behavior and how agents actually chose between different tools.
We're a team of researchers, engineers and work with the fastest growing AI first companies like Parallel, Firecrawl, Telnyx, TinyFish and more.
We're a team of researchers, engineers and theorists, in person in SF. You'll be building the benchmarks and the systems that run them.
Here are the broad themes that you’ll be working on
Benchmarks and evals design - designing domain-specific benchmarks from scratch: what to measure,...
Browse similar: AI jobs · AI startup jobs · Startup jobs · Founding team roles · San Francisco Bay Area