Backed by Lightspeed.
About the role
Judgment is the learning infrastructure for AI agents. Agents in production don't improve from prompts alone. They improve from experience: the tasks they attempt, the mistakes they make, the edge cases they hit. Here's how it works: We ingest everything your agents do in production: traces, tool calls, decisions, outcomes
More about this role
Judgment is the learning infrastructure for AI agents. Agents in production don't improve from prompts alone. They improve from experience: the tasks they attempt, the mistakes they make, the edge cases they hit. Here's how it works:
We ingest everything your agents do in production: traces, tool calls, decisions, outcomes
Judgment turns that raw experience into structured signals: failure modes, behaviors, rubrics, evals
Teams close the loop, shipping agent improvements validated against real production evidence
You'll own problems end-to-end: talking to customers, defining what to build, building it, and iterating until it's great. This is not a role where you implement specs handed down.
Investigation interfaces: Design how engineers understand what their systems did and why. Long traces, tool calls, decisions, failures. How do you make a complex sequence of events legible in minutes?
Verification: Build the platform for verifying system changes: hosted simulated environments, trajectory replay, and monitors for unintended behavior changes.
The improvement loop: Build the workflows that turn production data into datasets, evaluations, and regression checks, so the path from...
Browse similar: AI jobs · AI startup jobs · Startup jobs · San Francisco Bay Area