Environments and datasets for RL. Powering the world. Backed by Y Combinator.
About the role
Frontier labs test their models in environments like ours, which means we are the first to see novel model behavior. You will make sure that alignment is baked into everything the do: the code we deploy, the runbooks we write, and the papers we publish. Run experiments on how our environments shape behavior: reward hacking, cheating the task, unsafe use of tools, deception over long runs. Then change how we build them.
What they're looking for
- You are an engineer first. You build your own test pipelines, learn strange codebases quickly, and find the bug in the log
- You know LLM safety in depth: attacks, red teaming, monitoring and control evaluations, reward hacking
- You can turn a vague worry about a model into an experiment, run it, and say what the result does and does not prove
- You read transcripts closely. You catch the small thing that makes the whole result wrong
- You have put research into production systems, especially post-training, distillation, or evaluation that something depends on
- You write clearly enough that your results change what other people do
More about this role
Abundant is an applied research lab focused on scaling reinforcement learning for safe and reliable agentic capabilities. We are an extremely talent-dense team of researchers, roboticists, founders, and operators whose work includes the Waymo Driver.
Frontier labs test their models in environments like ours, which means we are the first to see novel model behavior. You will make sure that alignment is baked into everything the do: the code we deploy, the runbooks we write, and the papers we publish.
Run experiments on how our environments shape behavior: reward hacking, cheating the task, unsafe use of tools, deception over long runs. Then change how we build them.
Break our own graders. Find the ways to score well without doing the work, close them, and keep the attacks as tests everyone runs.
Build the tools that watch agents: monitors over trajectories, red team harnesses, and safety benchmarks that stay hard as models improve. On real runs, not toy ones.
Keep our environments sealed. Network isolation, escapes, credentials, and how much damage an agent can do when it turns on us. Last summer’s failures came from here.
Choose a research question about oversight, control, or...
Browse similar: AI jobs · AI startup jobs · Startup jobs · San Francisco Bay Area