# Senior Applied Scientist, Efficient LLM Inference & Model Optimization at Nebius

- Company: Nebius
- What the company does: Build and scale faster on the purpose-built AI cloud, engineered from silicon to API. Backed by Accel.
- Company website: https://nebius.com/
- Type: Startups (AI role)
- Level: Senior
- Location: Palo Alto, California, United States
- Work setup: On-site
- Posted: 2026-07-22
- Apply by: 2026-10-12
- Apply: https://careers.nebius.com/?gh_jid=4921549101
- Page: https://www.1752.vc/careers/jobs/nebius-senior-applied-scientist-efficient-llm-inference-and-model-optimization/

## About the role

Nebius Token Factory needs scientists who can turn frontier inference bottlenecks into research problems, publish credible work, and then help ship the results into production. This is not a papers-only research role. The Applied Scientist is expected to design rigorous experiments, write strong code, collaborate with engineers, and convert research into deployed inference capabilities.

## What they're looking for

- PhD in computer science, machine learning, ML systems, computer systems, computer architecture, electrical engineering, applied math, or a closely related field
- Strong publication record or equivalent research artifacts in ML, ML systems, efficient inference, model compression, quantization, distillation, serving systems, or related areas
- Strong hands-on coding ability in Python and PyTorch, ability to move from idea to experiment to prototype quickly
- Deep understanding of LLMs, VLMs, transformer inference, decoding algorithms, model compression, quantization, and production-serving tradeoffs
- Strong experimental design skills, including ablations, baselines, metrics, statistical reasoning, and failure analysis
- Excellent written and verbal communication

Tags: ML
