# Cyber Evaluations Engineer at Anthropic

- Company: Anthropic
- What the company does: Anthropic is an AI safety and research company that's working to build reliable, interpretable, and steerable AI systems. Backed by Accel, Bessemer and General Catalyst.
- Company website: https://www.anthropic.com/
- Type: Startups (AI role)
- Level: Mid level
- Location: Remote-Friendly, United States; San Francisco, CA | Washington, DC
- Work setup: Remote
- Posted: 2026-09-01
- Apply by: 2026-10-16
- Apply: https://job-boards.greenhouse.io/anthropic/jobs/5406367008
- Page: https://www.1752.vc/careers/jobs/anthropic-cyber-evaluations-engineer/

## About the role

We're hiring Cyber Evaluations Engineers to build and run the evaluations that measure cyber-relevant capabilities and safeguard robustness in our models. You'll design new evals, run per-release robustness testing, and dig into data on jailbreaks and prompt bypasses to understand where our safeguards hold up and where they don't. You'll also design many of the probes that detect cyber abuse in production and help shape the overall detection architecture alongside the policy team.

## What they're looking for

- Experience building or running evaluations, benchmarks, or test suites for software or ML systems, including delivering results on short, fixed timelines
- Hands-on cybersecurity experience (e.g., CTF participation, vulnerability research, exploit development, or security research)
- Proficiency in Python
- Strong ability to communicate evaluation results with multiple cross-functional stakeholders or potential policy stakeholders
- Deep offensive-security or security-research experience, including experience building AI security benchmarks
- Experience analyzing adversarial or abuse data (e.g., jailbreaks, prompt bypasses, intrusion or fraud telemetry)

Tags: Safeguards (Trust & Safety)
