Simulation environments to train & evaluate long-horizon AI agents. Backed by Y Combinator.
About the role
We’re hiring Software Engineers to build the simulation environments, tasks, and verifiers that challenge frontier models. You’ll help create the training and evaluation grounds that make it possible to measure and improve autonomous agents on realistic, challenging work.
More about this role
Polymath is an applied research lab focused on advancing long-horizon agent capabilities through reinforcement learning. We design and scale simulation environments where agents learn to operate safely and autonomously. We work with the world’s leading model labs to push the frontier of agent capabilities. We've raised >$8M from Base10, Founders Future, Y Combinator, and other incredible investors & angels.
We’re hiring Software Engineers to build the simulation environments, tasks, and verifiers that challenge frontier models. You’ll help create the training and evaluation grounds that make it possible to measure and improve autonomous agents on realistic, challenging work.
Building diverse, high-fidelity environments that test agents in realistic settings
Designing complex tasks that require long-horizon reasoning and tool use
Developing robust verifiers that reliably measure agent performance
Improving infrastructure and tooling to run, debug, and improve environments
Working closely with the research team to identify failure modes and turn them into new tasks and benchmarks
Have strong engineering fundamentals
Enjoy building from first principles and solving open-ended...
Browse similar: AI jobs · AI startup jobs · Startup jobs · Remote jobs · San Francisco Bay Area