# Cloud Inference Engineer at Luminal

- Company: Luminal (YC S25)
- What the company does: Making AI run fast on any hardware. Backed by Y Combinator and Felicis.
- Company website: https://luminal.com
- Type: Startups (AI role)
- Level: Mid level
- Location: San Francisco, CA, US
- Work setup: On-site
- Pay: $150K to $250K base salary per year (USD)
- Posted: 2025-10-16
- Apply by: 2026-10-08
- Apply: https://www.ycombinator.com/companies/luminal/jobs/J87U2SS-cloud-inference-engineer
- Page: https://www.1752.vc/careers/jobs/luminal-cloud-inference-engineer/

## About the role

Luminal (YC S25) builds an AI compiler and serving stack that makes models 10x faster and production ready with one line. Founding, on site in downtown SF. Ship low latency, high throughput model serving on Luminal Cloud.

## What they're looking for

- CUDA + GPU inference optimization
- vLLM, SGLang, or TensorRT-LLM experience
- KV caching, paged attention, batching, token streaming, etc
- Distributed compute (with GPUs is a super plus)
- No degree required

Tags: Engineering
