Startups · AI

Founding Machine Learning Infrastructure Engineer

Model AI · Palo Alto · On-site

← All jobs
About Model AI

ModelOp's Enterprise AI Command Center is the system of record that unifies every AI asset so you bring ML, GenAI, and agentic AI to production 10× faster. Backed by a16z.

About the role

We are looking for an ML Systems Engineer to help build and optimize the core serving infrastructure behind Agent Cloud. This role focuses on high-performance inference across different accelerators. You will work on model serving performance, accelerator utilization, long-context inference, batching, scheduling, KV cache management, runtime efficiency, and cost reduction. This is a deeply technical role at the intersection of ML systems, infrastructure, and product.

What they're looking for

  • Strong experience in ML systems, distributed systems, or high-performance computing
  • Experience optimizing inference or training workloads for large models
  • Familiarity with TPUs, GPUs, or other accelerators
  • Experience with one or more of CUDA, Triton, NCCL, JAX/XLA, PyTorch internals, vLLM, SGLang, TensorRT-LLM, distributed inference, or distributed training
  • Strong systems debugging skills
  • Comfort working across model code, runtime, infrastructure, and product requirements
More about this role

Peano AI is building the infrastructure and application stack for the next generation of agentic AI systems .

We believe token usage will grow exponentially over the coming years, but routing all inference through closed model providers will remain too expensive for many users and enterprises. Our thesis is that agentic applications require a vertically integrated stack: high-throughput, cost-efficient serving infrastructure paired with an application layer designed for long-running, agentic workloads.

Peano AI is building the Agent Cloud, a serving and training infrastructure platform purpose-built for agentic workloads, long-context inference, and large-scale open-source model deployment. By combining infrastructure and application design, we aim to make open-source models significantly more performant, practical, and competitive.

We are looking for an ML Systems Engineer to help build and optimize the core serving infrastructure behind Agent Cloud. This role focuses on high-performance inference across different accelerators.

You will work on model serving performance, accelerator utilization, long-context inference, batching, scheduling, KV cache management, runtime...

Read the full posting on Model AI's site ↗

Technical Staff

Build your edge while you search

Free tools for founders and investors, plus VC Unfiltered, our take on startups, venture and the people who build them.