Startups · AI

Research Engineer - Pre-training

Pluralis · San Francisco · Remote

← All jobs
About Pluralis

Pluralis Research works on Protocol Learning — decentralized, communication-efficient model-parallel training for foundation models. Backed by USV.

About the role

Distributed pretraining : Implement and optimize model-parallel training. Data, pipeline, and tensor parallelism for large models on heterogeneous GPUs under low-bandwidth, high-latency links. Performance optimization : Implement techniques that reduce communication overhead while maintaining model convergence in challenging network environments.

What they're looking for

  • Hands-on distributed training (required) : You've trained models across many devices in PyTorch with FSDP, DeepSpeed, Megatron, or your own implementation. You understand data, tensor, and pipeline parallelism
  • Strong engineering : Production-quality Python. Concurrency, failure handling, profiling before optimizing
  • Evidence of execution : Shipped systems, research code, open-source work, or serious personal projects
  • Mission alignment : You believe Protocol Learning is the viable third path for collective, trustless, and sovereign AI
More about this role

Pluralis Research works on Protocol Learning: training and serving large models in a fully decentralized way on small consumer-grade devices connected via the internet. Despite being dismissed as infeasible, we have made significant advances on this problem, most recently Agora, a permissionless run that pretrained an 8B model from scratch on consumer GPUs spread over the internet, with no single participant ever holding the full weights ( tech report ). While many of the core research problems have been solved, Protocol Learning unlocks a series of new challenges. For the mission in full, read A Third Path: Protocol Learning .

This setting breaks nearly every assumption of datacenter training: communication-efficient training across different parallelism axes, fault tolerance as nodes join and drop mid-run, heterogeneous compute and networks, and robustness to malicious participants. Our published methods include Subspace Networks , Factored Gossip DiLoCo , AsyncMesh , and Sentinel .

As a Research Engineer you'll build the training system that takes Protocol Learning from the 8B run to frontier scale: large models on heterogeneous hardware, in physically different regions,...

Read the full posting on Pluralis's site ↗

Research

Build your edge while you search

Free tools for founders and investors, plus VC Unfiltered, our take on startups, venture and the people who build them.