Vouch is the insurance broker that powers ambition. We’re a tech-enabled insurance advisory and brokerage purpose-built for growing companies in technology, life sciences, and professional services. Backed by Index, Y Combinator and 500 Global.
About the role
This is genuinely interesting work, and we say that well-aware of how every job posting claims that. We are betting on a specific approach to AI in a regulated domain, and we are not going to lay it all out in this posting — we'd rather it stay our edge until you're across the table from us. The hard problems live exactly where you'd hope — making an LLM system durable, auditable, and measurably improving, in a domain where being wrong has consequences.
What they're looking for
- Frontier models are instruments you play daily — in how you build (coding agents, multi-model workflows, letting an agent draft while you direct) and in what you build (planners, classifiers, tool-calling systems)
- You come at problems with approaches nobody asked for, and you are willing to be wrong out loud in service of moving the work
- "The model can probably do this" is a hypothesis you test with an eval, not a hope you ship on vibes
- You prototype fast, generalize what survives contact with reality, and delete what doesn't
- You ship in small, traceable increments, week after week, teammates can find the ticket from your branch name and the reasoning in your design doc
- Production is yours. You fix the OOM at the right layer and treat a correctness rework as finishing the job, not a "fast follow"
More about this role
The Role
Vouch is building AI software for judgment-heavy insurance work: a system that learns from experts and gets measurably better week over week. We are early — a small team, real experts, real production usage, real customers, and a lot of unanswered technical questions.
This is genuinely interesting work, and we say that well-aware of how every job posting claims that. We are betting on a specific approach to AI in a regulated domain, and we are not going to lay it all out in this posting — we'd rather it stay our edge until you're across the table from us. The hard problems live exactly where you'd hope — making an LLM system durable, auditable, and measurably improving, in a domain where being wrong has consequences.
You would join early in the system's life, in a rapidly evolving codebase that already carries more test code than source code. That ratio is not an accident; it is our style.
How we work, concretely: reasoning is written down and public — design docs land as pull requests, root-cause writeups happen in the channel, and demos are async videos every Friday morning. We ship to staging many times a day and to production behind consent-based pushes; standups are...
Browse similar: AI jobs · AI startup jobs · Startup jobs