Startups

Research Manager, Evaluation

Aaru · NYC · On-site

← All jobs
About Aaru

Humans are unreliable narrators of their own behavior. Memory fails, incentives distort answers, and social pressure warps what people say away from what they actually do. Backed by General Catalyst, Felicis and Redpoint.

About the role

As Evaluation Research Manager, you will lead a focused team of Evaluation Researchers and research engineers. You will translate Aaru's evaluation charter into a coherent portfolio of studies, datasets, and shared infrastructure, and you will be accountable for the quality, pace, and usefulness of the team's work.

What they're looking for

  • Two teams disagree about whether a system improved because they use different metrics and test sets. Identify the underlying construct, choose the right evidence, and create a shared evaluation that resolves the disagreement
  • A population looks plausible one profile at a time but fails to reproduce important real-world relationships. Build measurements that expose the gap and help Population Research identify its cause
  • A prediction method is well calibrated overall but systematically overconfident for a high-value subgroup. Determine whether the issue is data coverage, model structure, selection, condition shift, or the evaluation itself
  • An offline benchmark has become a development target and is beginning to leak into decisions. Redesign the evaluation system so teams can iterate quickly without exhausting the integrity of the final holdout
  • A component metric improves, but customer decisions do not. Determine whether the metric is invalid, the effect is too small, downstream components erase the gain, or the product is presenting the result incorrectly
  • Ground truth is delayed, noisy, incomplete, or open to multiple interpretations. Design a study that remains useful without pretending that the label is cleaner than it is
More about this role

Aaru builds simulations of human behavior. Each simulation contains a population of AI agents, each representing a person who could plausibly exist in the real world and capable of making decisions within a modeled environment. Companies and institutions use these simulations to test consequential choices before committing—from product launches and pricing decisions to strategic communications and policy changes.

Building a useful simulation requires more than generating plausible text. Populations must represent real people and groups; predictions must be calibrated; simulations must remain coherent as conditions change; and the product must make the resulting evidence legible enough to support real decisions.

We are a small, in-person team in New York. We work with urgency, high ownership, and intellectual honesty. We expect people to surface inconvenient evidence, change their minds quickly, and carry important work all the way to a result.

Evaluation Research determines whether Aaru's populations, predictions, and end-to-end simulations correspond closely enough to the real world to support consequential decisions. The team defines what should be measured, develops the methods...

Read the full posting on Aaru's site ↗

Technical Staff

Build your edge while you search

Free tools for founders and investors, plus VC Unfiltered, our take on startups, venture and the people who build them.