Terac is an AI‑native research platform that sources participants, conducts human‑like interviews at scale, analyzes results, and pays out participants - delivering actionable insights in hours, not weeks. Backed by Emergence Capital.
About the role
Review real user interaction traces with an AI shopping assistant Identify logical failures, inaccuracies, or poor recommendations in the text
More about this role
We're hiring AI evaluators to assess the accuracy and helpfulness of a new digital shopping assistant. This project focuses on understanding how well the system handles real-world e-commerce queries and where it falls short in its logic. Your analysis will directly feed into improving the underlying model and its response quality.
You will review real interaction traces between users and the shopping assistant within our custom platform. As you analyze these conversations, you will pinpoint specific failures, logical errors, or unhelpful product recommendations. From there, you will create structured rubrics and verifiers to consistently judge future response quality. This is an ongoing remote engagement requiring 20+ hours per week.
This opportunity is ideal for quality assurance specialists, AI data evaluators, and e-commerce professionals with a strong eye for detail. We welcome applicants with prior experience in prompt engineering, complex data annotation, or software testing. You should be comfortable analyzing text interactions deeply and building structured evaluation frameworks from scratch.
Review real user interaction traces with an AI shopping assistant
Identify...
Browse similar: AI jobs · AI startup jobs · Startup jobs · Remote jobs