Startups · AI

Senior Software Engineer, AI Evals

Sentry.co · San Francisco, California · Remote

← All jobs
About Sentry.co

Offline credential manager. Backed by Antler.

About the role

As a Senior Software Engineer on Sentry’s AI/ML team, you’ll be responsible for building the evaluation infrastructure that measures the accuracy, reliability, and real-world performance of our AI systems. This role is critical to ensuring that our debugging agents and AI-powered features behave correctly, safely, and predictably as they scale. You’ll design datasets, benchmarks, and test harnesses that turn ambiguous AI behavior into measurable signals, helping the team ship AI with confidence.

What they're looking for

  • Minimum 5+ years of professional experience with a Bachelor’s degree in computer science, machine learning, or a related field
  • Experience building testing, evaluation, or data infrastructure for complex systems (AI/ML experience strongly preferred)
  • Comfort writing production-quality code (we use Python and TypeScript)
  • Experience working with structured and unstructured datasets, labeling workflows, or data quality pipelines
  • Familiarity with modern ML systems and evaluation techniques (e.g., offline metrics, online evaluation, regression testing for models or prompts)
  • Bonus: experience evaluating LLMs, agentic systems, or AI-assisted developer tools
More about this role

Software runs the world and the pace is faster than ever. Sentry helps developers fix errors and performance issues before users notice, so teams can spend less time firefighting and more time building.

Trusted by 200,000+ organizations, Sentry is today’s application monitoring standard and our team is building its AI-native future.

As a Senior Software Engineer on Sentry’s AI/ML team, you’ll be responsible for building the evaluation infrastructure that measures the accuracy, reliability, and real-world performance of our AI systems. This role is critical to ensuring that our debugging agents and AI-powered features behave correctly, safely, and predictably as they scale. You’ll design datasets, benchmarks, and test harnesses that turn ambiguous AI behavior into measurable signals, helping the team ship AI with confidence.

Design and build robust evaluation frameworks to measure accuracy, reliability, regressions, and edge cases in AI systems

Create and curate high-quality datasets, golden test cases, and benchmarks grounded in real production data

Build automated test harnesses and metrics pipelines to continuously evaluate models, prompts, and agentic workflows

Partner closely...

Read the full posting on Sentry.co's site ↗

Engineering

Build your edge while you search

Free tools for founders and investors, plus VC Unfiltered, our take on startups, venture and the people who build them.