# Member of Technical Staff - Inference at Primeintellect

- Company: Primeintellect
- What the company does: Train, deploy, and continuously improve your own models on an integrated compute, training, inference, and sandbox stack. Backed by Menlo.
- Company website: https://www.primeintellect.ai/
- Type: Startups (AI role)
- Level: Senior
- Location: Remote
- Work setup: Remote
- Posted: 2026-07-08
- Apply by: 2026-10-08
- Apply: https://jobs.ashbyhq.com/PrimeIntellect/abfa70f7-a6f1-44d2-a6c1-560e1c8477d4
- Page: https://www.1752.vc/careers/jobs/primeintellect-member-of-technical-staff-inference/

## About the role

This is a hybrid position spanning cloud LLM serving, LLM inference optimization and RL systems. You will be working on advancing our ability to evaluate and serve models trained with our RL Lab at scale. The two key areas are: You'll join a team of experienced engineers and researchers working on cutting-edge problems in AI infrastructure. We believe in open development and encourage team members to contribute to the broader AI community through research and open-source contributions.

## What they're looking for

- Building ML Systems at Scale: 3+ years building and running large‑scale ML/LLM services with clear latency/availability SLOs
- Inference Backends: Hands‑on with at least one of vLLM, SGLang, TensorRT‑LLM
- Distributed Serving Infra: Familiarity with distributed and disaggregated serving infrastructure such as NVIDIA Dynamo
- Inference Internals: Deep understanding of prefill vs. decode, KV‑cache behavior, batching, sampling, speculative decoding, parallelism strategies
- Full‑Stack Debugging: Comfortable debugging CUDA/NCCL, drivers/kernels, containers, service mesh/networking, and storage, owning incidents end‑to‑end
- Python: Systems tooling and backend services

Tags: Engineering
