Private, domain-specific benchmarks in legal, tax, and finance. Backed by a16z and Pear VC.
About the role
You, alongside our team, will own the platform that runs our benchmarks. This spans everything needed to evaluate LLMs at scale: Python libraries, a web platform, distributed systems, cloud infrastructure, and tooling. You'll work across the stack—whatever needs to be built to run benchmarks reliably and efficiently.
What they're looking for
- Technical
- Strong engineering fundamentals : You can build and ship quickly with high quality. You should have a track record of building things of significant scope (within jobs, side projects, open source, etc.)
- Python expertise : Significant experience in Python, especially in a professional setting
- System Design Experience: You should be familiar with common concepts like VMs, containerization, load balancers, databases, etc. and when to use them appropriately
- Familiarity with LLMs: You should have experience working with LLM APIs, and understand concepts like temperature, tokenization, reasoning, etc
- Learning velocity: The role encompasses a wide variety of tasks. Rather than expecting you to be an expert on Day 1, we are looking for someone who can learn new skills and technologies quickly
More about this role
We are looking for exceptional engineers to join our team.
You, alongside our team, will own the platform that runs our benchmarks. This spans everything needed to evaluate LLMs at scale: Python libraries, a web platform, distributed systems, cloud infrastructure, and tooling. You'll work across the stack—whatever needs to be built to run benchmarks reliably and efficiently.
At Vals, we believe in autonomy. You will be given a high degree of independence to make decisions on tech stacks, system architecture, and code structure. You will also provide guidance to others on the team, both through informal feedback and formal processes like architecture and code reviews.
Our platform serves startups, enterprises, and research labs measuring model performance. We work with all the major foundation model labs, and some of the largest financial institutions and hospital systems in the world. Our work has been featured by the Wall Street Journal, New York Times, Washington Post, and Bloomberg.
We are building the standard for evaluating the ability of LLMs to perform real-world tasks. You will contribute directly to the infrastructure that makes this possible.
Build distributed systems to...
Browse similar: AI jobs · AI startup jobs · Startup jobs · San Francisco Bay Area