Startups · AI

Research Fellowship - Mechanistic Interpretability

Vmax · San Francisco · On-site

← All jobs
About Vmax

Vmax is automating reinforcement learning. We transform proprietary data and evals into new sets of environments. We refine agents on new examples of the tasks they are intended to perform. Backed by South Park Commons.

About the role

LLMs are fantastically powerful and there is a rapidly growing corpus of work devoted to understanding their internal representations and computations. We use the tools of mechanistic interpretability to enhance reinforcement learning by generating intrinsic rewards as a supplement or alternative to downstream human-generated verifiers.

What they're looking for

  • Track record of research excellence or strong research promise, demonstrated through publications, preprints, open-source work, technical projects, competitions, or publicly available artifacts
  • Working understanding of reinforcement learning
  • Familiarity with mechanistic interpretability, representation analysis, or empirical methods for understanding neural networks
  • Strong programming ability in Python and experience with at least one major ML framework such as PyTorch or JAX
  • Clear written and verbal communication of technical ideas
More about this role

V max is an applied research lab developing AI capable of open-ended learning. We are building systems to exceed humans in all capacities by optimizing beyond the local maxima of learning from human expertise.

LLMs are fantastically powerful and there is a rapidly growing corpus of work devoted to understanding their internal representations and computations. We use the tools of mechanistic interpretability to enhance reinforcement learning by generating intrinsic rewards as a supplement or alternative to downstream human-generated verifiers.

This 3 to 6 month fellowship is for PhD students or equivalent early-career researchers who want to work at the intersection of mechanistic interpretability and reinforcement learning. You will own a focused research project, work closely with Vmax technical staff, and contribute to research publications.

  • Develop mechanistic interpretability methods for understanding internal representations, features, circuits, and computations in language models and agents.
  • Investigate how model internals can be used to generate intrinsic rewards, auxiliary objectives, diagnostics, or training signals for reinforcement learning.
  • Design and run...

Read the full posting on Vmax's site ↗

Research and Engineering

Build your edge while you search

Free tools for founders and investors, plus VC Unfiltered, our take on startups, venture and the people who build them.