# Research Engineer - Distributed Training at Primeintellect

- Company: Primeintellect
- What the company does: Train, deploy, and continuously improve your own models on an integrated compute, training, inference, and sandbox stack. Backed by Menlo.
- Company website: https://www.primeintellect.ai/
- Type: Startups (AI role)
- Level: Mid level
- Location: San Francisco
- Work setup: On-site
- Posted: 2026-07-08
- Apply by: 2026-10-08
- Apply: https://jobs.ashbyhq.com/PrimeIntellect/8bd52610-175c-42a7-a7cd-b29c45f9d305
- Page: https://www.1752.vc/careers/jobs/primeintellect-research-engineer-distributed-training/

## About the role

Build and optimize the distributed training infrastructure behind our pre-training and large-scale RL training workloads by contributing to our prime-rl framework. Improve end-to-end training efficiency across compute, memory, networking, and scheduling layers.

## What they're looking for

- Strong systems engineering experience in AI/ML infrastructure, especially around large-scale model training or inference
- Deep familiarity with PyTorch and distributed training frameworks such as PyTorch Distributed, DeepSpeed, FSDP, Megatron, vLLM, Ray, or related tooling
- Experience optimizing training performance across kernels, memory movement, communication overhead, or parallelization strategy
- Hands-on experience with large-scale training techniques including data parallelism, tensor parallelism, and pipeline parallelism
- Strong understanding of GPU architecture, profiling, and performance debugging
- Ability to identify bottlenecks across the stack and drive improvements from first principles

Tags: Research
