Vmax is automating reinforcement learning. We transform proprietary data and evals into new sets of environments. We refine agents on new examples of the tasks they are intended to perform. Backed by South Park Commons.
About the role
LLMs are fantastically powerful and there is a rapidly growing corpus of work devoted to understanding their internal representations and computations. We use the tools of mechanistic interpretability to enhance reinforcement learning by generating intrinsic rewards as a supplement or alternative to downstream human-generated verifiers.
What they're looking for
- Track record of research excellence or strong research promise, demonstrated through publications, preprints, open-source work, technical projects, competitions, or publicly available artifacts
- Working understanding of reinforcement learning
- Familiarity with mechanistic interpretability, representation analysis, or empirical methods for understanding neural networks
- Strong programming ability in Python and experience with at least one major ML framework such as PyTorch or JAX
- Clear written and verbal communication of technical ideas
More about this role
V max is an applied research lab developing AI capable of open-ended learning. We are building systems to exceed humans in all capacities by optimizing beyond the local maxima of learning from human expertise.
LLMs are fantastically powerful and there is a rapidly growing corpus of work devoted to understanding their internal representations and computations. We use the tools of mechanistic interpretability to enhance reinforcement learning by generating intrinsic rewards as a supplement or alternative to downstream human-generated verifiers.
This 3 to 6 month fellowship is for PhD students or equivalent early-career researchers who want to work at the intersection of mechanistic interpretability and reinforcement learning. You will own a focused research project, work closely with Vmax technical staff, and contribute to research publications.
- Develop mechanistic interpretability methods for understanding internal representations, features, circuits, and computations in language models and agents.
- Investigate how model internals can be used to generate intrinsic rewards, auxiliary objectives, diagnostics, or training signals for reinforcement learning.
- Design and run...
Browse similar: AI jobs · AI startup jobs · Startup jobs · San Francisco Bay Area