# Researcher, Evaluations and Benchmarks at Alice

- Company: Alice
- What the company does: Secure the AI apps and agents you build and deploy with red teaming and runtime guardrails powered by real-world adversarial data. Backed by Norwest and CRV.
- Company website: https://www.alice.io/
- Type: Startups (AI role)
- Level: Mid level
- Location: Remote, United States
- Work setup: Remote
- Posted: 2026-10-05
- Apply by: 2026-11-19
- Apply: https://alice.io/positions/position-d9_27d?utm_source=1752vc&utm_medium=careers
- Page: https://www.1752.vc/careers/jobs/alice-researcher-evaluations-and-benchmarks/

## About the role

You ship a benchmark every two to three weeks ( example benchmark ). Each one measures a frontier risk that nobody has measured yet. Some go public. Some go only to the labs. You will not write every eval yourself. Each benchmark pairs you with an in-house researcher who owns that harm area, and you get a budget for freelancers you direct. You own the taxonomy, the harness, the quality bar and the release. Found on 1752vc Careers, the job board for startup and VC roles.

## What they're looking for

- PhD or Masters in computer science, machine learning or a related field, or equivalent depth from industry research
- 3+ years building and running safety or security evaluations for language models in production, at an AI lab, a model provider, or a safety and security research organisation
- 5+ relevant research publications in the field of AI safety and security including lead author on at least 2 of them
- Strong engineer. Evaluation harnesses, distributed inference, vLLM, reading a codebase and fixing it
- You can build a taxonomy, not only score against one
- You can direct a researcher and two freelancers without managing them formally

Tags: R&D

Source: 1752vc Careers, https://www.1752.vc/careers/jobs/alice-researcher-evaluations-and-benchmarks/
