# Member of Technical Staff (Language Model Evaluations) at Artificial Analysis

- Company: Artificial Analysis
- What the company does: Comparison and analysis of AI models and API hosting providers. Independent benchmarks across key performance metrics including quality, price, output speed & latency. Backed by AI Grant.
- Company website: https://artificialanalysis.ai/
- Type: Startups (AI role)
- Level: Senior
- Location: San Francisco (On-site)
- Work setup: On-site
- Posted: 2026-07-29
- Apply by: 2026-10-08
- Apply: https://jobs.ashbyhq.com/artificialanalysis/2b4df1cc-2a51-4516-8b7f-2994ef3e5616
- Page: https://www.1752.vc/careers/jobs/artificial-analysis-member-of-technical-staff-language-model-evaluations/

## About the role

Language model evaluation is the sharpest question in AI: what can these systems actually do? Our answers, from the Artificial Analysis Intelligence Index to AA-Omniscience, AA-Briefcase and our coding agent evaluations, are the reference the industry uses. We’re hiring Members of Technical Staff to build the next generation of them.

## What they're looking for

- You have deep, hands-on experience evaluating language models and strong opinions about why most benchmarks fail
- • 3+ years of relevant professional experience, across industry or research
- • Strong analytical and critical thinking skills
- • Strong Python, with hands-on experience running evaluation harnesses and building datasets
- • Deep familiarity with the LLM evaluation landscape: the major benchmarks and their failure modes, contamination, preference-based methods, and agentic evaluation
- • Strong statistical grounding: you know when a result is signal and when it is noise

Tags: Technical Staff (incl. Product)
