Startups · AI

Software Engineer, Evaluation Platform / Infra

Thinking Machines · San Francisco · On-site

← All jobs
About Thinking Machines

Connectionism: Research Blog by Thinking Machines Lab. Backed by a16z, Accel and GV.

About the role

Evaluation is one of the most important pillars of building frontier AI systems. It guides research direction, powers experimentation, and helps us understand whether changes to data and training are improving the capabilities and behaviors we care about.

What they're looking for

  • A bachelor’s degree, or equivalent practical experience, in computer science, engineering, machine learning, or a related field
  • Two years of post-grad work experience as a software engineer or ML engineer, exclusive of internships
  • Hands-on experience building or maintaining evaluations, benchmarks, graders, or model-quality systems for large language or multimodal models
  • Strong software engineering fundamentals and experience building reliable, maintainable systems
  • Proficiency in at least one backend programming language, we primarily use Python and Rust. We use React and Typescript on the frontend
  • Experience with databases, data pipelines, distributed systems, or other data-intensive infrastructure
More about this role

The mission of Thinking Machines is to build AI that extends human will and judgment. We are training frontier models with Inkling, developing Tinker to let people make models their own, and crafting interfaces that broaden human-AI communication. We believe the future worth building is human, and we're hiring people who want to build it.

Evaluation is one of the most important pillars of building frontier AI systems. It guides research direction, powers experimentation, and helps us understand whether changes to data and training are improving the capabilities and behaviors we care about.

To support this work, researchers need a powerful, self-serve platform that makes it easy to author evaluations, run them or reproduce them reliably at scale, and extract insight from the results. The platform must support both standardized external benchmarks and fast-moving internal evaluations, many kinds of tasks and graders, and inspection from aggregate metrics down to individual model trajectories.

In this role, you will design and build this platform end to end. You will work across Python frameworks, data pipelines, APIs, and user-facing applications, and collaborate closely with...

Read the full posting on Thinking Machines's site ↗

Core Engineering

Build your edge while you search

Free tools for founders and investors, plus VC Unfiltered, our take on startups, venture and the people who build them.