Terac is an AI‑native research platform that sources participants, conducts human‑like interviews at scale, analyzes results, and pays out participants - delivering actionable insights in hours, not weeks. Backed by Emergence Capital.
About the role
Pick a familiar workflow and convert it into a demanding AI prompt Test your prompt in ChatGPT to find failure points, making it harder if the AI succeeds
More about this role
We're running a paid study to build a bench of people who are exceptionally good at designing tasks that expose AI model limitations. By turning real-world workflows into demanding requests, we can better evaluate where current models break down. This initial trial helps us identify individuals suited for ongoing prompt engineering and evaluation work.
You will spend about an hour translating a complex workflow from your job or personal life into a demanding prompt that requires reasoning and real-world lookup. After running it in ChatGPT to identify where the model fails, you will refine the prompt until it breaks the system. Finally, you will write a clear grading rubric that a stranger could use to evaluate any AI's attempt at your task. This entire process is screen-recorded, as we are assessing your thought process just as much as the final submitted files.
We welcome professionals, domain experts, and power users who have deep knowledge of specific workflows. You need to be capable of evaluating an AI's output within seconds and comfortable working on a laptop or desktop with a ChatGPT account. Candidates who excel at this trial will be considered for a long-term bench of...
Browse similar: AI jobs · AI startup jobs · Startup jobs · Remote jobs