Startups · AI

Machine Learning Engineer - ML Training Platform

Pluralis · San Francisco · Remote

← All jobs
About Pluralis

Pluralis Research works on Protocol Learning — decentralized, communication-efficient model-parallel training for foundation models. Backed by USV.

About the role

Multi-cloud infrastructure : Design the resource management systems that provision and orchestrate compute across AWS, GCP, and Azure with infrastructure-as-code (Pulumi/Terraform). Handle dynamic scaling, state synchronization, and concurrent operations across hundreds of heterogeneous nodes.

What they're looking for

  • Infrastructure and platform engineering (required) : Production experience with infrastructure-as-code (Pulumi/Terraform/CloudFormation) managing multi-cloud deployments, Docker/Kubernetes (EKS), GPU workloads, and heterogeneous clusters at scale
  • Distributed systems and ML infrastructure : You understand distributed training workflows: checkpointing, data sharding, model versioning, long-running job orchestration
  • Decentralized networking : P2P, NAT traversal, traffic shaping, real bandwidth constraints
  • Systems programming and reliability : Strong Python engineering (asyncio, concurrency, retry logic, cloud SDKs, CLI tooling) with hands-on observability and SRE practice, Prometheus/Grafana, performance profiling, incident response
  • Environment fit : You've done this in a startup with heavy service orchestration, or at big-tech scale, and you can show which systems you owned
  • Mission alignment : You believe Protocol Learning is the viable third path for collective, trustless, and sovereign AI
More about this role

Pluralis Research works on Protocol Learning: training and serving large models in a fully decentralized way on small consumer-grade devices connected via the internet. Despite being dismissed as infeasible, we have made significant advances on this problem, most recently Agora, a permissionless run that pretrained an 8B model from scratch on consumer GPUs spread over the internet, with no single participant ever holding the full weights ( tech report ). While many of the core research problems have been solved, Protocol Learning unlocks a series of new challenges. For the mission in full, read A Third Path: Protocol Learning .

Our training and inference doesn't happen in a datacenter. It happens on consumer nodes and cloud instances that are not co-located, connected by ordinary internet, joining and leaving mid-run. Your primary role is to architect, build, and scale the platform that keeps continuous experimentation and large-scale training running on top of that: infrastructure orchestration, distributed compute, and the services that tie them together.

Multi-cloud infrastructure : Design the resource management systems that provision and orchestrate compute across AWS, GCP,...

Read the full posting on Pluralis's site ↗

Engineering

Build your edge while you search

Free tools for founders and investors, plus VC Unfiltered, our take on startups, venture and the people who build them.