# Senior/Staff System Research Engineer – LLM Inference Optimization at Snowflake

- Company: Snowflake
- What the company does: Snowflake powers AI, data engineering, applications, and analytics on a trusted, scalable AI Data Cloud—eliminating silos and accelerating innovation. Backed by Sequoia.
- Company website: https://www.snowflake.net/
- Type: Startups (AI role)
- Level: Senior
- Location: US-WA-Bellevue
- Work setup: Remote
- Pay: $236K to $330K base salary per year (USD)
- Posted: 2026-09-21
- Apply by: 2026-11-05
- Apply: https://jobs.ashbyhq.com/snowflake/9e0ae021-f264-4efb-bc80-26c04bd06df6
- Page: https://www.1752.vc/careers/jobs/snowflake-senior-staff-system-research-engineer-llm-inference-optimization/

## About the role

Design and develop high-performance LLM inference systems, spanning distributed serving, runtime systems, GPU execution, and performance-critical kernels. Develop novel techniques to improve inference latency, generation speed, throughput, memory efficiency, scalability, and cost.

## What they're looking for

- Bachelor’s degree in Computer Science, Electrical Engineering, or a related field. A Master’s degree or PhD is preferred
- 5+ years of experience in one or more of the following areas: LLM inference systems, distributed AI systems, GPU systems, or high-performance computing
- Strong understanding of modern LLM inference architectures and the performance tradeoffs involved in serving large-scale models
- Hands-on experience with modern LLM inference and serving frameworks, such as vLLM, SGLang, TensorRT-LLM, or similar systems
- Experience designing, extending, or optimizing inference runtimes, including areas such as scheduling, batching, KV-cache management, distributed execution, parallelism, speculative decoding, or disaggregated serving
- Strong understanding of GPU architectures and experience with CUDA, Triton, or similar GPU programming environments

Tags: Engineering
