Mercor is organizing human intelligence to power the AI economy. We are powering frontier research, AI benchmarks, and AI agent training at scale for the top AI labs and enterprises. Backed by General Catalyst and Menlo.
About the role
Enterprise agents are complex systems, and they only pay off when their work is reliable and economically viable. Evaluation is how you get both: checking correctness is the obvious case, and routing is the subtler one, since choosing a model against cost, latency, and quality requires quality to be measurable at all. Knowing where the bar sits is the hard part. You decompose real work, take the standard from the practitioners who hold it, and encode it so an agent cannot shortcut it.
What they're looking for
- Professional, academic, or research experience in agent engineering and evaluation, including how agent runtimes and harnesses produce a trajectory and where it fails
- Experience building evaluation suites for LLM or agent systems, and familiarity with how benchmarks such as terminal-bench, tau-bench, and APEX are constructed and where they get gamed
- Judgment about task and rubric design: turning a fuzzy notion of quality into something measurable, with agent or model improvements to show for it
- Strong software engineering fundamentals, and the ability to work independently on ambiguous, loosely specified problems
- Bonus: experience with Harbor environments and RL environments
More about this role
Mercor's mission is to organize human intelligence to power the AI economy. We're a leading AI data company, building the layer between human expertise and frontier models. Millions of domain experts on the platform are paid over $4 million per day to train frontier AI models. Mercor's APEX benchmark family measures AI's real-world impact on professional work. Mercor Enterprise brings this same infrastructure to Fortune 500 companies: helping companies capture how their best people actually work, translating that expertise directly back into agents.
Mercor is creating a new category of work where expertise powers AI advancement. Achieving this requires an ambitious, fast-paced and deeply committed team. You’ll work alongside researchers, operators, and AI companies at the forefront of shaping the systems that are redefining society. Mercor is a profitable Series C company valued at $10 billion. We work in-person five days a week in our San Francisco, NYC, or London offices.
Enterprise agents are complex systems, and they only pay off when their work is reliable and economically viable. Evaluation is how you get both: checking correctness is the obvious case, and routing is the...
Browse similar: Startup jobs · San Francisco Bay Area