Startups · AI

Product engineer, full stack

Judgment Labs · San Francisco · On-site

← All jobs
About Judgment Labs

Backed by Lightspeed.

About the role

Judgment is the learning infrastructure for AI agents. Agents in production don't improve from prompts alone. They improve from experience: the tasks they attempt, the mistakes they make, the edge cases they hit. Here's how it works: We ingest everything your agents do in production: traces, tool calls, decisions, outcomes

More about this role

Judgment is the learning infrastructure for AI agents. Agents in production don't improve from prompts alone. They improve from experience: the tasks they attempt, the mistakes they make, the edge cases they hit. Here's how it works:

We ingest everything your agents do in production: traces, tool calls, decisions, outcomes

Judgment turns that raw experience into structured signals: failure modes, behaviors, rubrics, evals

Teams close the loop, shipping agent improvements validated against real production evidence

You'll own problems end-to-end: talking to customers, defining what to build, building it, and iterating until it's great. This is not a role where you implement specs handed down.

Investigation interfaces: Design how engineers understand what their systems did and why. Long traces, tool calls, decisions, failures. How do you make a complex sequence of events legible in minutes?

Verification: Build the platform for verifying system changes: hosted simulated environments, trajectory replay, and monitors for unintended behavior changes.

The improvement loop: Build the workflows that turn production data into datasets, evaluations, and regression checks, so the path from...

Read the full posting on Judgment Labs's site ↗

Engineering

Build your edge while you search

Free tools for founders and investors, plus VC Unfiltered, our take on startups, venture and the people who build them.