# Machine Learning Research Scientist, Evaluations at Scale

- Company: Scale
- What the company does: Backed by Accel, Index and Y Combinator.
- Company website: http://www.scale.com
- Type: Startups (AI role)
- Level: Mid level
- Location: San Francisco, CA; Seattle, WA; New York, NY
- Work setup: On-site
- Posted: 2026-08-26
- Apply by: 2026-10-10
- Apply: https://job-boards.greenhouse.io/scaleai/jobs/4728014005
- Page: https://www.1752.vc/careers/jobs/scale-machine-learning-research-scientist-evaluations/

## About the role

Scale works with the industry's leading AI labs to provide high quality data and accelerate progress in GenAI research. We are looking for Research Scientists and Research Engineers with expertise in LLM post-training (SFT, RLHF, reward modeling) and evaluation. This role is on the evaluation pod within the GenAI Research Organization and will focus on building benchmarks and diagnosing model failure modes in both text and multimodal modalities.

## What they're looking for

- Design and build benchmarks and evaluation methods that measure LLM capabilities in both text and multimodal modalities
- Apply post-training expertise (SFT, RLHF, reward modeling) to connect observed failures to the data and training interventions that address them
- Publish research findings in top-tier AI conferences
- Ph.D. or Master's degree in Computer Science, Machine Learning, AI, or a related field
- Deep understanding of deep learning, reinforcement learning, and large-scale model fine-tuning
- Experience with post-training techniques such as RLHF, preference modeling, or instruction tuning, and with LLM evaluation or benchmark development

Tags: Research
