# Machine Learning Engineer - ML Training Platform at Pluralis

- Company: Pluralis
- What the company does: Pluralis Research works on Protocol Learning — decentralized, communication-efficient model-parallel training for foundation models. Backed by USV.
- Company website: https://pluralis.ai/
- Type: Startups (AI role)
- Level: Mid level
- Location: San Francisco
- Work setup: Remote
- Posted: 2026-08-31
- Apply by: 2026-10-15
- Apply: https://jobs.ashbyhq.com/pluralis-research/7b107585-5a3b-428d-ad5d-1c85be9dca4e
- Page: https://www.1752.vc/careers/jobs/pluralis-machine-learning-engineer-ml-training-platform/

## About the role

Multi-cloud infrastructure : Design the resource management systems that provision and orchestrate compute across AWS, GCP, and Azure with infrastructure-as-code (Pulumi/Terraform). Handle dynamic scaling, state synchronization, and concurrent operations across hundreds of heterogeneous nodes.

## What they're looking for

- Infrastructure and platform engineering (required) : Production experience with infrastructure-as-code (Pulumi/Terraform/CloudFormation) managing multi-cloud deployments, Docker/Kubernetes (EKS), GPU workloads, and heterogeneous clusters at scale
- Distributed systems and ML infrastructure : You understand distributed training workflows: checkpointing, data sharding, model versioning, long-running job orchestration
- Decentralized networking : P2P, NAT traversal, traffic shaping, real bandwidth constraints
- Systems programming and reliability : Strong Python engineering (asyncio, concurrency, retry logic, cloud SDKs, CLI tooling) with hands-on observability and SRE practice, Prometheus/Grafana, performance profiling, incident response
- Environment fit : You've done this in a startup with heavy service orchestration, or at big-tech scale, and you can show which systems you owned
- Mission alignment : You believe Protocol Learning is the viable third path for collective, trustless, and sovereign AI

Tags: Engineering
