Backed by a16z.
About the role
The interesting problem is not getting an agent to test a control once. It is knowing whether it was right across hundreds of controls, for customers whose methodologies all differ, where the ground truth is a senior auditor's judgment rather than a label in a dataset. Building the harness and the evaluation that make that tractable is the core of this role, and it would be yours to own.
What they're looking for
- Experience building on LLMs and agentic systems: harnesses, long-horizon agents, evals, or resilient workflows that survive contact with messy real-world input
- Experience building enterprise SaaS products that handle sensitive customer data
- Comfort working across the stack: solid backend fundamentals (APIs, data modeling, distributed systems) plus enough frontend fluency to ship a feature end to end
- Experience with AWS, Kubernetes, and infrastructure as code
- Experience in fintech, audit, governance, risk, compliance, or another regulated environment
- A track record of high ownership at an early- or growth-stage startup, where you shaped not just the code but the product and the process around it
More about this role
Petual is the AI-powered control tester for the modern enterprise. Our agents test whether a company's financial, operational, and technology controls actually work, and trace every conclusion to its source.
The hard part isn't speed. It's being right. External auditors re-perform our work, so precision on sample 1 has to equal precision on sample 400, and an agent that can't support a conclusion has to say so rather than guess. We're backed by Andreessen Horowitz; the team comes from Stripe, Retool, and Lyft; and public companies already run their SOX programs on Petual.
The interesting problem is not getting an agent to test a control once. It is knowing whether it was right across hundreds of controls, for customers whose methodologies all differ, where the ground truth is a senior auditor's judgment rather than a label in a dataset. Building the harness and the evaluation that make that tractable is the core of this role, and it would be yours to own.
Own the agent harness at the core of the product: the workflows, checkpoints, and evaluation that keep long-horizon agents reliable across hours of testing
Design and own the backend services and APIs that carry evidence from...
Browse similar: AI jobs · AI startup jobs · Startup jobs · San Francisco Bay Area