# Distributed Training and Inference Engineer at Sciforium

- Company: Sciforium
- What the company does: Sciforium builds the next generation of AI models with unprecedented efficiency, privacy, and versatility. Backed by SignalFire.
- Company website: https://sciforium.com
- Type: Startups (AI role)
- Level: Mid level
- Location: San Francisco, CA
- Work setup: On-site
- Pay: $190K to $250K base salary per year (USD)
- Posted: 2026-08-24
- Apply by: 2026-10-14
- Apply: https://jobs.ashbyhq.com/Sciforium/1471adc1-cf58-4174-8dc2-e8c0e29cfb17
- Page: https://www.1752.vc/careers/jobs/sciforium-distributed-training-and-inference-engineer/

## About the role

Sciforium is seeking a highly skilled Distributed Training and Inference Engineer to build, optimize, and maintain the critical software stack that powers our large-scale AI training and serving workloads. In this role, you will work across the entire machine learning infrastructure from low-level CUDA/ROCm runtimes to high-level frameworks like JAX and PyTorch to ensure our distributed training systems are fast, scalable, stable, and efficient.

## What they're looking for

- 5+ years of industry experience in ML systems, distributed training, or related fields
- Bachelor’s or Master’s degree in Computer Science, Computer Engineering, Electrical Engineering, or related technical fields
- Strong programming experience in Python, C++, and familiarity with ML tooling and distributed systems
- Deep understanding of profiling tools (e.g., Nsight, ROCm Profiler, XLA profiler, TPU tools)
- Deep expertise with partitioning configuration on the modern ML frameworks such as PyTorch and JAX
- Experience with multi-node distributed training systems and orchestration frameworks (DTensor, GSPMD, etc.)

Tags: Engineering
