Chat, compare, vote for the world's best AI models. Join the community shaping the public leaderboard for LLMs, image, and code models through real-world evaluation. Backed by Lightspeed and a16z.
About the role
This is an engineering role. The work is in the product code: the authorization and identity layer, the API surfaces our users and partners depend on, and the libraries and patterns other engineers build on top of. You will design systems, write production code, and ship on the same cadence as the rest of engineering.
What they're looking for
- We care most about two things: recent hands-on production engineering, and code-level application security judgment
- Application and API security: you apply it at the code level, not the checklist level. Authorization models, session and token handling, and multi-tenant isolation, with a specific problem you found or fixed that you can explain down to the mechanism
- 6+ years of software or security engineering experience, with meaningful time building and shipping production systems at scale
- Strong proficiency in a modern backend language, and the judgment to design interfaces other engineers will live with for years
- Solid data fundamentals. You're comfortable modeling and querying Postgres, you know where security decisions belong in a data layer, and you have run migrations against production data without breaking it
- Working fluency with cryptographic application primitives: HMACs, secure random generation, key and secret rotation, and the ways sensitive data leaks through logs, traces, and stored payloads
More about this role
Arena is the platform for evaluating how AI models perform in the real world. Founded by researchers from UC Berkeley's SkyLab, we're on a mission to measure and advance the frontier of AI for real-world use, and to build the foundation for everyone to understand, shape, and benefit from it.
Tens of millions of people use Arena each month to evaluate how frontier systems handle the work they actually do. The preferences they share power the most transparent, rigorous, and human-centered evaluations in AI. Leading AI labs, enterprises, and independent researchers rely on our work and open datasets to understand how models behave in real workflows: agentic coding, creative generation, professional productivity, and beyond. We go beyond leaderboards and decompose what human experience reveals about AI, so models advance toward the work people actually do.
We're a team of researchers, academics, builders, and creatives from UC Berkeley, Google, Stanford, and DeepMind. We seek truth, move fast, and value craftsmanship, curiosity, and impact over hierarchy. We're building a company where thoughtful, curious people from all backgrounds can do their best work together, in an office...
Browse similar: AI jobs · AI startup jobs · Startup jobs · San Francisco Bay Area