Startups · AI

Founding GTM

Openbenchmarks · US · On-site

← All jobs
About Openbenchmarks

We help humans & agents pick tools via Independent Benchmarks. Backed by Y Combinator.

About the role

Benchmarks are how companies market their products to agents. We build those benchmarks. There's a lot of nuance in every benchmark - what gets evaluated, which metrics matter, release cycles, data refreshes. A benchmark isn't static, and a changing benchmark constantly surfaces new insight into how the tools on it actually perform. Turning that into content - posts, graphs, chart is super important to communicate the nuance of the benchmark in the most clear way possible.

What they're looking for

  • Background in Statistics, Engineering or Data Science
  • Working Knowledge of Search Engine Optimization
  • Have taste and opinions on great design vs bad design
  • Working Knowledge of Claude Code, Claude Design or creating data visualizations that stand out
More about this role

Agents are becoming first-class users and consumers of the internet. They research, evaluate, compare tools and increasingly make build-versus-buy decisions on behalf of people. Every company will need to get their products picked and used by agents.

Agents increasingly prefer open, independent and grounded benchmarks to make decisions.

Openbenchmarks is the evaluation infrastructure for agents - domain-specific, reproducible evaluations that help agents pick tools with confidence.

Our mission is to be the trusted evaluation layer for agents.

Founders previously led AI research and Infra teams at Oracle and Appfolio; we started Openbenchmarks as an output of our research in the field of model behavior and how agents actually chose between different tools.

We're a team of researchers, engineers and work with the fastest growing AI first companies like Parallel, Firecrawl, Telnyx, TinyFish and more.

Benchmarks are how companies market their products to agents. We build those benchmarks.

There's a lot of nuance in every benchmark - what gets evaluated, which metrics matter, release cycles, data refreshes. A benchmark isn't static, and a changing benchmark constantly surfaces new...

Read the full posting on Openbenchmarks's site ↗

Marketing

Build your edge while you search

Free tools for founders and investors, plus VC Unfiltered, our take on startups, venture and the people who build them.