Startups

Member of Technical Staff - Research Engineer

Black Forest Labs · San Francisco (United States) · On-site

← All jobs
About Black Forest Labs

Backed by a16z.

About the role

Improve the performance, reliability, and numerical stability of production training runs for large multimodal generative models Profile full training steps across model code, attention, kernels, data loading, encoders, communication, optimizer steps, checkpointing, and memory pressure

What they're looking for

  • Experience working deeply on large-scale training systems, ideally as part of a training group working closely with researchers
  • Strong PyTorch fluency, including comfort reading and modifying low-level training code rather than only using high-level APIs
  • Experience with distributed training concepts such as FSDP, tensor/model/context/sequence parallelism, activation checkpointing, NCCL, and overlapping compute and communication
  • Hands-on experience improving training throughput, memory footprint, or stability in real training runs
  • Experience profiling GPU workloads with tools like Nsight Systems, Nsight Compute, torch profiler, trace viewers, or custom telemetry
  • Practical GPU performance judgment: you may use modern coding agents and tools as much as you want, but you need the understanding to verify correctness, numerical behavior, and performance, and to own the result
More about this role

We're the team behind Latent Diffusion, Stable Diffusion, and FLUX—foundational technologies that changed how the world creates images and video. We’re creating the generative models that power how people make images and video—tools used by millions of creators, developers, and businesses worldwide. Our FLUX models are among the most advanced in the world, and we're just getting started.

Headquartered in Freiburg, Germany with a growing presence in San Francisco, we’re scaling fast while staying true to what makes us different: research excellence, open science, and building technology that expands human creativity.

Large-scale training is where research ideas become real, and where many of the hardest problems are no longer cleanly separated into “research” or “engineering.” A promising architecture only matters if we can train it stably, efficiently, and correctly across large GPU fleets.

In this role, you will be embedded in production training and help where the hardest systems and performance problems arise: attention performance, custom kernels, low-precision training, profiling, memory behavior, data movement, distributed training stability, and throughput regressions. You...

Read the full posting on Black Forest Labs's site ↗

Research

Build your edge while you search

Free tools for founders and investors, plus VC Unfiltered, our take on startups, venture and the people who build them.