World Labs is a spatial intelligence company, building frontier models that can perceive, generate, and interact with the 3D world. Backed by NEA and a16z.
About the role
We are looking for a Performance Engineer to make World Labs’ models train and serve as fast as the hardware allows.
What they're looking for
- You should excel at the fundamentals below — we index on inference, serving, GPU optimization, and training performance. Distributed-systems breadth is welcome, but secondary
- Strong performance-engineering foundations: profiling, roofline analysis, latency/throughput optimization, and disciplined root-cause investigation
- Deep GPU programming and optimization experience (CUDA and/or Triton) — kernel-level tuning, memory hierarchy, and bandwidth optimization at scale
- Hands-on experience optimizing inference and serving for large models: batching, KV/prompt caching, quantization, and low-latency, high-throughput sampling
- Hands-on experience optimizing training performance: parallelism, distributed communication, mixed/low precision, and utilization
- Working knowledge of ML framework internals (PyTorch and/or JAX, torch.compile, XLA, or similar compiler paths)
More about this role
World Labs is a frontier AI research and product company advancing spatial intelligence, the next frontier beyond large language models. Co-founded by Dr. Fei-Fei Li , Justin Johnson and Ben Mildenhall , the company is pioneering world models that perceive, generate, reason, and interact with virtual and physical worlds.
The company’s flagship product, Marble , transforms text, images, and video into fully navigable 3D worlds, unlocking applications across gaming, film, architecture, robotics, and immersive digital experiences. Backed by leading investors and with over $1B raised, World Labs is assembling a world-class team at the intersection of AI research and real-world deployment.
We are looking for a Performance Engineer to make World Labs’ models train and serve as fast as the hardware allows.
Running large generative world models at scale is a novel systems problem. You will find the bottlenecks — in kernels, in the serving path, in the training loop, in how we use our GPUs — and eliminate them. Your ownership is technical and concrete: the throughput you unlock, the latency you cut, the utilization you win back, and the correctness you hold while doing it. You will work up...
Browse similar: AI jobs · AI startup jobs · Startup jobs · San Francisco Bay Area