Startups · AI

Member of Technical Staff - Research

Vals AI · San Francisco · On-site

← All jobs
About Vals AI

Private, domain-specific benchmarks in legal, tax, and finance. Backed by a16z and Pear VC.

About the role

We are looking for exceptional researchers and research engineers to design and build the next generation of AI benchmarks. You will create high-impact, challenging evaluations that push the boundaries of what we can measure in foundation models. This role is ideal for someone with deep research expertise who wants to see their work directly influence how the world evaluates AI systems.

What they're looking for

  • Advanced research experience : Master's degree or PhD in Computer Science, NLP, Machine Learning, or related field. Undergrads with very strong research backgrounds may also be considered
  • Publication track record : Published papers in reputable venues (NeurIPS, ICML, ACL, EMNLP, etc.) with focus on NLP, ML evaluation, or benchmarking
  • Research methodology : Strong understanding of experimental design, statistical analysis, and evaluation frameworks
  • Technical skills : Proficiency in Python for research and experimentation
  • Communication : Ability to clearly communicate complex research ideas to both technical and non-technical audiences
  • Collaboration : Experience working in research teams and integrating feedback
More about this role

We are looking for exceptional researchers and research engineers to design and build the next generation of AI benchmarks. You will create high-impact, challenging evaluations that push the boundaries of what we can measure in foundation models. This role is ideal for someone with deep research expertise who wants to see their work directly influence how the world evaluates AI systems.

You will lead the design and development of novel benchmarks that assess real-world capabilities of LLMs. Our benchmarks shape how foundation models are developed and generative AI applications are built. We work with every major foundation model lab - along with leading financial institutions and the application-layer companies pushing the frontier forward. Our work has been featured by the Wall Street Journal, New York Times, Washington Post, and Bloomberg.

We are building the standard for evaluating the ability of LLMs to perform real-world tasks. You will be at the forefront of defining what that standard looks like.

Design and develop novel, high-impact benchmarks that assess challenging real-world capabilities

Conduct research to ensure our benchmarks are valid, reliable, and...

Read the full posting on Vals AI's site ↗

Engineering & Research

Build your edge while you search

Free tools for founders and investors, plus VC Unfiltered, our take on startups, venture and the people who build them.