Startups · AI

Research Operations, Reinforcement Learning

Anthropic · San Francisco, CA · On-site

← All jobs
About Anthropic

Anthropic is an AI safety and research company that's working to build reliable, interpretable, and steerable AI systems. Backed by Accel, Bessemer and General Catalyst.

About the role

Frontier research runs on more than research alone. Behind every model launch, review, and cross-team program is a web of decisions, priorities, and follow-through that keeps hundreds of researchers moving in the same direction. Research Operations owns that work: we bring clarity to ambiguity, provide stability during change, and continually deepen our domain knowledge so that researchers can spend their time on the problems only they can solve. It is one of the highest-leverage functions at Anthropic.

What they're looking for

  • Care deeply about Anthropic's mission and let it underpin all judgment calls you make across the org
  • Clear writing: you turn dense, jargon-heavy material into summaries and plans that technical collaborators across the RL org and beyond can act on
  • An eye for great dashboards: you can look at data visualizations and scope improvements that make lessons at-a-glance clearer and grounded
  • Strong follow-through and attention to detail, nothing falls through the cracks on your watch
  • Ability to push back on senior stakeholders to offer strategic guidance as needed, disagree without causing discord, and influence without direct authority
  • Comfort with ambiguity and a fast-changing environment, you stay steady when the pace is relentless and the stakes are as high as they can possibly be
More about this role

Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems.

Frontier research runs on more than research alone. Behind every model launch, review, and cross-team program is a web of decisions, priorities, and follow-through that keeps hundreds of researchers moving in the same direction. Research Operations owns that work: we bring clarity to ambiguity, provide stability during change, and continually deepen our domain knowledge so that researchers can spend their time on the problems only they can solve. It is one of the highest-leverage functions at Anthropic.

We are hiring an experienced operator to embed with the leadership of our reinforcement learning (RL) organization, the team responsible for training Claude to be more capable, more reliable, and more aligned. You will shape what leadership spends its time on, make sure priority decisions reach the right decision-makers and get made quickly, and clear the path so...

Read the full posting on Anthropic's site ↗

AI Research & Engineering

Build your edge while you search

Free tools for founders and investors, plus VC Unfiltered, our take on startups, venture and the people who build them.