Anthropic is an AI safety and research company that's working to build reliable, interpretable, and steerable AI systems. Backed by Accel, Bessemer and General Catalyst.
About the role
Anthropic's Safeguards organization builds the policies, evaluations, and detection and enforcement systems that define and hold the limits on how Claude can be used. In this role, you'll lead our policy design team, managing the teams responsible for radicalization, child safety, user well-being, harmful manipulation, and election integrity, among other harm areas. Found on 1752vc Careers, the job board for startup and VC roles.
What they're looking for
- Experience leading teams — including managing managers or senior specialists — in AI safety, product policy, or a related field
- Deep, applied familiarity with consumer harm areas such as child safety, mental health and well-being, manipulation, or election integrity, and good judgment about how these harms differ in mechanism, severity, and mitigation
- A track record of exceptional cross-team collaboration: building durable working relationships with teams you don't control, and getting to shared decisions where ownership is genuinely distributed
- Experience translating policy positions into mechanisms that can be enforced and measured, and communicating the reasoning to technical and non-technical audiences, including executives
- Sound judgment in ambiguous, high-consequence decisions, and comfort making a call and escalating appropriately on incomplete information
- Subject-matter depth in one or more of the portfolio's harm areas, from academia, clinical practice, civil society, government, or trust & safety work
More about this role
Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems.
Anthropic's Safeguards organization builds the policies, evaluations, and detection and enforcement systems that define and hold the limits on how Claude can be used. In this role, you'll lead our policy design team, managing the teams responsible for radicalization, child safety, user well-being, harmful manipulation, and election integrity, among other harm areas.
The team is responsible for understanding and defining the risks that come with engaging with Claude, how those risks materialize in the real world, and the mitigations needed to prevent them. As the manager, you'll work with your team to draw the boundaries between what is and is not allowed, then partner with research, product, and engineering to build the right interventions. Mitigating these harms takes the whole stack: the values and judgment trained into the model itself, the policies and...
Browse similar: AI jobs · AI startup jobs · Startup jobs · San Francisco Bay Area