Startups · AI

Cloud Inference Engineer

Luminal · San Francisco, CA, US · On-site

← All jobs
About Luminal

Making AI run fast on any hardware. Backed by Y Combinator and Felicis.

About the role

Luminal (YC S25) builds an AI compiler and serving stack that makes models 10x faster and production ready with one line. Founding, on site in downtown SF. Ship low latency, high throughput model serving on Luminal Cloud.

What they're looking for

  • CUDA + GPU inference optimization
  • vLLM, SGLang, or TensorRT-LLM experience
  • KV caching, paged attention, batching, token streaming, etc
  • Distributed compute (with GPUs is a super plus)
  • No degree required
More about this role
  • CUDA + GPU inference optimization
  • vLLM, SGLang, or TensorRT-LLM experience
  • KV caching, paged attention, batching, token streaming, etc.
  • Distributed compute (with GPUs is a super plus)
  • No degree required

Luminal (YC S25) builds an AI compiler and serving stack that makes models 10x faster and production ready with one line.

Founding, on site in downtown SF. Ship low latency, high throughput model serving on Luminal Cloud.

  • Deploy and tune models with optimizations like KV caching, paged attention, sequence packing, etc.
  • Conducting model performance reviews
  • Improve scheduler, batcher, autoscaling; profile latency, cost, utilization
  • Sometimes write kernels and, yes, occasional tasteful shitposting

Read the full posting on Luminal's site ↗

Engineering

Build your edge while you search

Free tools for founders and investors, plus VC Unfiltered, our take on startups, venture and the people who build them.