# Staff Research Engineer, Model Efficiency at Cohere

- Company: Cohere
- What the company does: Backed by AI Grant.
- Company website: https://cohere.ai/
- Type: Startups (AI role)
- Level: Senior
- Location: New York
- Work setup: Remote
- Posted: 2025-11-07
- Apply by: 2026-10-08
- Apply: https://jobs.ashbyhq.com/cohere/c80f0fe9-3fc4-49fe-9f26-a7115350b1fc
- Page: https://www.1752.vc/careers/jobs/cohere-staff-research-engineer-model-efficiency/

## About the role

Large Language Models (LLMs) continue to push the boundaries of what AI systems can do — but inference is still the bottleneck. The Model Efficiency team is responsible for pushing the limits of LLM inference efficiency across our foundation models. We explore and ship breakthroughs across the model execution stack, including:

## What they're looking for

- Have a PhD in Machine Learning or a related field
- Understand LLM architecture, and how to optimize LLM inference given resource constraints
- Have significant experience with one or more techniques that enhance model efficiency
- Strong software engineering skills
- An appetite to work in a fast-paced high-ambiguity start-up environment
- Publications at top-tier conferences and venues (ICLR, ACL, NeurIPS)

Tags: Modeling
