# Performance Engineer, Inference Engine at Anthropic

- Company: Anthropic
- What the company does: Anthropic is an AI safety and research company that's working to build reliable, interpretable, and steerable AI systems. Backed by Accel, Bessemer and General Catalyst.
- Company website: https://www.anthropic.com/
- Type: Startups (AI role)
- Level: Mid level
- Location: San Francisco, CA | New York City, NY
- Work setup: On-site
- Posted: 2026-09-09
- Apply by: 2026-10-24
- Apply: https://job-boards.greenhouse.io/anthropic/jobs/5418323008
- Page: https://www.1752.vc/careers/jobs/anthropic-performance-engineer-inference-engine/

## About the role

Anthropic's inference engine is the software between the accelerator kernels and the routing layer. It manages the entire token path in between: batching requests, laying the model out across chips, managing memory for weights and activations, coordinating every forward pass, and managing model state across requests. Built in-house, it runs on all of our accelerator platforms, serving Claude to millions of users and running our research workloads.

## What they're looking for

- A working mental model of LLM inference: how prefill and decode land on an accelerator's compute, memory, and interconnect, and what the host is doing meanwhile
- Proven quick learner: ramped fast in deep, unfamiliar systems and shipped consequential changes quickly
- Strong systems programming (Rust, C++, or similar), with care for code quality and tests
- Analytical about performance: observe and profile first, form a hypothesis, test it, then change the code and measure again
- Low ego: ask the naive question, take feedback well, pick up slack outside your job description
- Enjoy pair programming (we love to pair!) and care about the societal impacts of your work

Tags: AI Research & Engineering
