Backed by Lightspeed.
About the role
Judgment is the learning infrastructure for AI agents. Agents in production don't improve from prompts alone. They improve from experience: the tasks they attempt, the mistakes they make, the edge cases they hit. Here's how it works: We ingest everything your agents do in production: traces, tool calls, decisions, outcomes
More about this role
Judgment is the learning infrastructure for AI agents. Agents in production don't improve from prompts alone. They improve from experience: the tasks they attempt, the mistakes they make, the edge cases they hit. Here's how it works:
We ingest everything your agents do in production: traces, tool calls, decisions, outcomes
Judgment turns that raw experience into structured signals: failure modes, behaviors, rubrics, evals
Teams close the loop, shipping agent improvements validated against real production evidence
You'll build the product experiences that make this loop legible, and you'll build the agents that run it. This is not a role where you implement specs handed down. You'll own problems end-to-end: talking to customers, defining what to build, building it, and iterating until it's great.
Judgment Agent: Shape how the Judgment Agent runs large-scale investigations: parallel investigators working across thousands of production traces, each covering a different dimension (failure modes, tool errors, regressions, drift), merging results into one answer.
Verification: Build the platform for verifying agent changes: hosted simulated environments for stateful agent evals,...
Browse similar: AI jobs · AI startup jobs · Startup jobs · San Francisco Bay Area