Startups · AI

Performance Engineer (Inference, Training & GPU)

World Labs · San Francisco · On-site

← All jobs
About World Labs

World Labs is a spatial intelligence company, building frontier models that can perceive, generate, and interact with the 3D world. Backed by NEA and a16z.

About the role

We are looking for a Performance Engineer to make World Labs’ models train and serve as fast as the hardware allows.

What they're looking for

  • You should excel at the fundamentals below — we index on inference, serving, GPU optimization, and training performance. Distributed-systems breadth is welcome, but secondary
  • Strong performance-engineering foundations: profiling, roofline analysis, latency/throughput optimization, and disciplined root-cause investigation
  • Deep GPU programming and optimization experience (CUDA and/or Triton) — kernel-level tuning, memory hierarchy, and bandwidth optimization at scale
  • Hands-on experience optimizing inference and serving for large models: batching, KV/prompt caching, quantization, and low-latency, high-throughput sampling
  • Hands-on experience optimizing training performance: parallelism, distributed communication, mixed/low precision, and utilization
  • Working knowledge of ML framework internals (PyTorch and/or JAX, torch.compile, XLA, or similar compiler paths)
More about this role

World Labs is a frontier AI research and product company advancing spatial intelligence, the next frontier beyond large language models. Co-founded by Dr. Fei-Fei Li , Justin Johnson and Ben Mildenhall , the company is pioneering world models that perceive, generate, reason, and interact with virtual and physical worlds.

The company’s flagship product, Marble , transforms text, images, and video into fully navigable 3D worlds, unlocking applications across gaming, film, architecture, robotics, and immersive digital experiences. Backed by leading investors and with over $1B raised, World Labs is assembling a world-class team at the intersection of AI research and real-world deployment.

We are looking for a Performance Engineer to make World Labs’ models train and serve as fast as the hardware allows.

Running large generative world models at scale is a novel systems problem. You will find the bottlenecks — in kernels, in the serving path, in the training loop, in how we use our GPUs — and eliminate them. Your ownership is technical and concrete: the throughput you unlock, the latency you cut, the utilization you win back, and the correctness you hold while doing it. You will work up...

Read the full posting on World Labs's site ↗

Engineering

Build your edge while you search

Free tools for founders and investors, plus VC Unfiltered, our take on startups, venture and the people who build them.