A modern messaging platform for teams of humans and agents. Backed by Accel and Index.
About the role
We have real, longitudinal, multi-party workspace data, and human / agent users whose behavior tells you whether the systems you're designing are actually meaningfully improving. Working with live business communication means permissions, redaction, and security are things we have to think about as well. If you want to have your research come into contact with reality, Ando is the place for it.
What they're looking for
- Strong applied research background , with depth in model evaluation, benchmarking, and/or failure analysis. You've built evals you trusted enough to make decisions with
- Evidence over credentials. Work samples or code that demonstrate the skills: eval frameworks, benchmark suites, failure-analysis reports or tooling, labeling infrastructure. Show us something you built to find out whether a system actually worked
- Strong technical communication. You can explain complex ideas simply and hold high-bandwidth, generative technical conversations with researchers and with our product team
- Comfort with mess. Real workspace data is incomplete, ambiguous, and full of edge cases that break clean abstractions. You treat that as signal, not noise
More about this role
Ando is a messaging platform where AI agents take on work alongside their human teammates. We’re rebuilding Slack from the ground up around two core ideas: durable memory and agents as first-class participants.
We have real, longitudinal, multi-party workspace data, and human / agent users whose behavior tells you whether the systems you're designing are actually meaningfully improving. Working with live business communication means permissions, redaction, and security are things we have to think about as well. If you want to have your research come into contact with reality, Ando is the place for it.
Proactivity: An agent embedded in a team's channels has to decide, message by message, whether to ignore, quietly track, or intervene. Evaluating that judgment means building benchmarks where the ground truth includes silence . Existing agent benchmarks are almost entirely reactive.
Memory: What should a workspace agent remember across weeks and months of participation, in what representation, and how do you measure whether memory is helping versus hurting agent usefulness?
Continual Learning: How much does basic memory, retrieval, and context impact agent performance, vs where do we...
Browse similar: AI jobs · AI startup jobs · Startup jobs · San Francisco Bay Area