Startups · AI

ML Systems Engineer

ChipAgents · San Jose · On-site

← All jobs
About ChipAgents

AI-driven chip design and verification with ChipAgents. Iterate on Your Chip Design & Verification 10x Faster by Collaborating with ChipAgents in Your Favorite Code Editor. Backed by Bessemer.

About the role

We are seeking an ML Systems Engineer to optimize the performance and efficiency of large language model inference powering our agentic AI platform. This is a technical role focused on low-level systems optimization. You will implement performance optimizations, build evaluation harnesses, and architect multi-node clusters for training and inference that push the limits of LLM throughput and latency.

What they're looking for

  • B.S., M.S., or PhD in Computer Science, Electrical Engineering, or related field (or equivalent experience)
  • Experience with large-scale ML systems, GPU computing, or high-performance inference optimization
  • Strong proficiency in Python and C++/CUDA, hands-on experience with SGLang, vLLM, PyTorch, or similar inference frameworks
  • Deep understanding of GPU architecture, memory hierarchies, and parallel computing paradigms
  • Experience deploying and optimizing LLMs in production: model serving, batching strategies, distributed inference, or quantization
  • Strong systems-level debugging and profiling skills, comfort working at multiple layers of the stack from CUDA kernels to application logic
More about this role

ChipAgents is redefining the future of chip design and verification with agentic AI workflows. Our platform leverages cutting-edge generative AI to assist engineers in RTL design, simulation, and verification, dramatically accelerating chip development. Founded by experts in AI and semiconductor engineering, we partner with top semiconductor firms, cloud providers, and innovative startups to build intelligent AI agents. The company is a Series A company backed by tier-1 VC firms. ChipAgents is deployed in production to companies that have shipped 16B chips.

We are seeking an ML Systems Engineer to optimize the performance and efficiency of large language model inference powering our agentic AI platform. This is a technical role focused on low-level systems optimization. You will implement performance optimizations, build evaluation harnesses, and architect multi-node clusters for training and inference that push the limits of LLM throughput and latency. Your work will directly impact the responsiveness and cost-efficiency of AI agents used by leading semiconductor companies to design chips.

Design, deploy, and optimize LLM inference systems across multi-node clusters, maximizing...

Read the full posting on ChipAgents's site ↗

Engineering

Build your edge while you search

Free tools for founders and investors, plus VC Unfiltered, our take on startups, venture and the people who build them.