Join the team redefining how the world experiences design. Hey, g'day, mabuhay, kia ora,你好, hallo, vítejte! Thanks for stopping by. Backed by Bessemer, General Catalyst and 500 Global.
About the role
Canva's generative models are judged by millions of people who will never read a benchmark. They just know whether the design looks right. Turning that judgement into something measurable is the hardest problem in our research stack, and it gates everything else. If we cannot measure design quality reliably, we cannot train against it, we cannot tell a real improvement from noise, and we cannot decide what ships.
What they're looking for
- Experience linking evaluation metrics to downstream business or user outcomes, and diagnosing why they diverge
- Experience turning subjective human judgement into reliable, objective evaluation signal through rubric design, human data pipelines and model training
- Strong grounding in multimodal generative models (diffusion, transformers, VLMs and MLLMs) and their architectures, deep enough to know where evaluation will break
- Experience with reward modelling, preference learning, or alignment methods involving human feedback
- Experience setting technical direction across multiple teams or a wide specialty area in a globally distributed organisation, with the instinct to find the gap nobody owns and close it without waiting for a mandate
More about this role
Canva's generative models are judged by millions of people who will never read a benchmark. They just know whether the design looks right. Turning that judgement into something measurable is the hardest problem in our research stack, and it gates everything else. If we cannot measure design quality reliably, we cannot train against it, we cannot tell a real improvement from noise, and we cannot decide what ships.
We are looking for a Principal Research Scientist who defines what evaluation needs to become as the space gets harder, rather than running the playbook we already have. You will own how Canva evaluates generative quality across the whole of Canva Research, including problems we have not framed yet: new modalities, evaluation that reflects real differences between content types, user segments and markets, and a much tighter link between what our metrics say and what users and the business actually experience.
This is a Canva-wide craft leadership role, setting direction across our research groups in Australia, Europe, the US and China. You will be the person others come to when the numbers and the eyes disagree.
The evaluation strategy for Canva Research . Define what...
Browse similar: AI jobs · AI startup jobs · Startup jobs · San Francisco Bay Area