dmodel: steerable and explainable AI. Backed by Y Combinator.
About the role
As a Member of Technical Staff, you'll build the reinforcement learning environments and evaluations we use to study how AI agents approach alignment problems. This role is a strong fit for someone early in their research career who wants hands-on experience in AI safety, interpretability, and reinforcement learning. Build reinforcement learning environments and evals to test how AI agents approach alignment problems
More about this role
d_model is a fundamental AI research lab partnering with frontier labs to turn their models into capable interpretability and alignment researchers. Alongside our partnerships, we aim to use the agents we build for independent research.
Our team brings experience from places including OpenAI, Google, Anthropic, EleutherAI, and MATS.
As a Member of Technical Staff, you'll build the reinforcement learning environments and evaluations we use to study how AI agents approach alignment problems. This role is a strong fit for someone early in their research career who wants hands-on experience in AI safety, interpretability, and reinforcement learning.
Build reinforcement learning environments and evals to test how AI agents approach alignment problems
Develop graders robust to specification gaming
Find novel interpretability and alignment techniques
Present findings in team research meetings and contribute to papers in ML
Run experiments and identify patterns in model behavior and failure modes
Investigate AI Interpretability techniques within our reinforcement learning environments
Have strong proficiency in Python and ML frameworks (PyTorch or JAX)
Are able to iterate quickly and...
Browse similar: AI jobs · AI startup jobs · Startup jobs · San Francisco Bay Area