# Member of Technical Staff, Model Efficiency at Cohere

- Company: Cohere
- What the company does: Backed by AI Grant.
- Company website: https://cohere.ai/
- Type: Startups (AI role)
- Level: Senior
- Location: New York
- Work setup: Remote
- Posted: 2025-11-07
- Apply by: 2026-10-08
- Apply: https://jobs.ashbyhq.com/cohere/2a989030-6d14-4924-88c1-d878911e26fa
- Page: https://www.1752.vc/careers/jobs/cohere-member-of-technical-staff-model-efficiency/

## About the role

Our team is a fast-growing group of researchers and engineers focused on building reliable ML systems and pushing the boundaries of LLM inference efficiency. We develop techniques that improve how models execute in production, driving lower latency, higher throughput, and consistent quality across diverse workloads.

## What they're looking for

- 5+ years of experience writing high-performance, production-quality code
- Strong programming skills in C++ or Python (Rust/Go also welcome)
- Experience working with large language models and familiarity with the LLM inference ecosystem (e.g., vLLM, SGLang, etc.)
- Ability to diagnose and resolve performance bottlenecks across the model execution stack
- A strong bias for action — you ship fast, measure impact, and iterate
- GPU programming, CUDA, or low-level systems optimization

Tags: Modeling
