# Senior Staff LLM Inference Engineer at d-Matrix

- Company: d-Matrix
- What the company does: d-Matrix is redefining AI inference with memory-centric compute built for ultra-low latency, greater efficiency, and scalable AI infrastructure.
- Company website: https://www.d-matrix.ai
- Type: Startups (AI role)
- Level: Senior
- Location: Santa Clara
- Work setup: Remote
- Pay: $195K to $285K base salary per year (USD)
- Posted: 2026-08-18
- Apply by: 2026-10-08
- Apply: https://jobs.ashbyhq.com/d-Matrix/f36b38d8-1d7c-40ad-9649-0c9bdef8429a
- Page: https://www.1752.vc/careers/jobs/d-matrix-senior-staff-llm-inference-engineer/

## About the role

Identify and prototype emerging LLM inference use cases suited to heterogeneous hardware deployments. Build compelling proof-of-concept systems that demonstrate D-Matrix capabilities to customers, partners, and internal stakeholders.

## What they're looking for

- Bachelor’s degree in Computer Science, Electrical Engineering, or a related field, and 10+ years of relevant engineering experience or equivalent demonstrated experience
- Master’s or PhD in Computer Science, Electrical Engineering, or a related field preferred, with 6+ years of relevant industry experience
- Strong proficiency in Python and C/C++
- Hands-on experience optimizing LLM inference — attention kernels, KV cache, batching strategies, quantization (INT8/FP8/INT4)
- Experience with at least one major inference framework (vLLM, SGLang, TensorRT-LLM, ONNX Runtime, or similar) at a contributor level
- Familiarity with GPU kernel programming (CUDA/Triton) and performance profiling tools

Tags: Architecture
