Startups · AI

Senior AI/ML Engineer — LLM & Agent Stack (Customer Facing)

TrueFoundry · San Mateo, San Francisco Bay Area · On-site

← All jobs
About TrueFoundry

Deploy, secure, and scale LLMs and AI agents with TrueFoundry's enterprise AI Gateway and MCP Gateway. Rate limiting, observability, and compliance built in. Backed by Peak XV.

About the role

Architect and implement scalable agent orchestration patterns (graph-based executors, state management, multi-agent coordination) for production workloads. Own critical integrations: model adapters, LLM gateway hooks, vector DBs, tools & external APIs, and the platform’s LLMops flows.

What they're looking for

  • 5–10 years of software engineering with substantial experience building distributed systems, infra, or ML platforms
  • Deep practical experience integrating and deploying LLMs in production (RAG, retrieval, embeddings pipelines)
  • Hands-on experience with agent orchestration frameworks (LangGraph / LangChain or custom agent runtimes) and stateful workflow design
  • Strong systems knowledge: Kubernetes, container orchestration, service meshes, and performance tuning
  • Proven track record building observability, cost controls, and policy enforcement for production services
More about this role

About TrueFoundry

Every production AI system, whether it's powering customer support, writing code, analyzing financial data, or diagnosing medical conditions, needs the same foundational infrastructure.A way to route between models. A way to manage tools and integrate them securely. A way to orchestrate agents and enforce governance. A unified compute layer to run it all.

We are looking for a Senior AI/ML Engineer — LLM & Agent Stack (Customer-Facing) to join the team.

Companies are moving beyond simple chatbots to production agentic systems. These systems route between OpenAI, Anthropic, Google, and self-hosted models. They integrate dozens of tools via protocols like MCP. They orchestrate multi-agent workflows where agents coordinate with other agents.

The infrastructure to support this doesn't exist yet. You can't just duct-tape together a few API calls and call it production-ready.

  • Intelligent routing with observability, cost policies, and fallback logic
  • Centralized tool and MCP server management with security and lifecycle controls
  • Agent orchestration with governance and guardrails
  • A unified compute layer to run self-hosted models, custom tools, and agents

AI Gateway...

Read the full posting on TrueFoundry's site ↗

Engineering

Build your edge while you search

Free tools for founders and investors, plus VC Unfiltered, our take on startups, venture and the people who build them.