# Performance Engineer (Inference, Training & GPU) at World Labs

- Company: World Labs
- What the company does: World Labs is a spatial intelligence company, building frontier models that can perceive, generate, and interact with the 3D world. Backed by NEA and a16z.
- Company website: https://www.worldlabs.ai/
- Type: Startups (AI role)
- Level: Mid level
- Location: San Francisco
- Work setup: On-site
- Pay: $200K to $300K base salary per year (USD)
- Posted: 2026-05-01
- Apply by: 2026-10-08
- Apply: https://jobs.ashbyhq.com/worldlabs/207875a2-d8d9-4b9c-ba9f-a19358c07575
- Page: https://www.1752.vc/careers/jobs/world-labs-performance-engineer-inference-training-and-gpu/

## About the role

We are looking for a Performance Engineer to make World Labs’ models train and serve as fast as the hardware allows.

## What they're looking for

- You should excel at the fundamentals below — we index on inference, serving, GPU optimization, and training performance. Distributed-systems breadth is welcome, but secondary
- Strong performance-engineering foundations: profiling, roofline analysis, latency/throughput optimization, and disciplined root-cause investigation
- Deep GPU programming and optimization experience (CUDA and/or Triton) — kernel-level tuning, memory hierarchy, and bandwidth optimization at scale
- Hands-on experience optimizing inference and serving for large models: batching, KV/prompt caching, quantization, and low-latency, high-throughput sampling
- Hands-on experience optimizing training performance: parallelism, distributed communication, mixed/low precision, and utilization
- Working knowledge of ML framework internals (PyTorch and/or JAX, torch.compile, XLA, or similar compiler paths)

Tags: Engineering
