Startups · AI

Research Engineer - Post-Training

Pluralis · San Francisco · Remote

← All jobs
About Pluralis

Pluralis Research works on Protocol Learning — decentralized, communication-efficient model-parallel training for foundation models. Backed by USV.

About the role

Build the post-training stack : You build the RL training loop end-to-end: rollout ingestion from the geo-distributed inference pipeline, reward computation, policy updates, and getting updated weights back out to the network. You set the direction, and you make things happen.

What they're looking for

  • Strong engineering : Production-quality Python and PyTorch: concurrency, failure handling, profiling before optimizing
  • Research ability : Publications in RL post-training, asynchronous or distributed RL, or nearby fields are a strong signal. So is unpublished work you can defend in detail
  • Mission alignment : You believe Protocol Learning is the viable third path for collective, trustless, and sovereign AI
More about this role

Pluralis Research works on Protocol Learning: training and serving large models in a fully decentralized way on small consumer-grade devices connected via the internet. Despite being dismissed as infeasible, we have made significant advances on this problem, most recently Agora, a permissionless run that pretrained an 8B model from scratch on consumer GPUs spread over the internet, with no single participant ever holding the full weights ( tech report ). While many of the core research problems have been solved, Protocol Learning unlocks a series of new challenges. For the mission in full, read A Third Path: Protocol Learning .

Agora gave us a pretrained 8B model. Post-training is how we make it useful for agentic use-cases. But every post-training stack you've seen assumes a datacenter — synchronous rollouts, fast interconnects, trusted workers. Ours gets none of that. It has to run on consumer GPUs, and Macs spread across the public internet, training a model whose weights no single participant ever holds, with rollouts arriving from a geo-distributed inference pipeline at high latencies. Your primary role is to make RL post-training work here anyway — the algorithms and the...

Read the full posting on Pluralis's site ↗

Research

Build your edge while you search

Free tools for founders and investors, plus VC Unfiltered, our take on startups, venture and the people who build them.