Humans are unreliable narrators of their own behavior. Memory fails, incentives distort answers, and social pressure warps what people say away from what they actually do. Backed by General Catalyst, Felicis and Redpoint.
About the role
As Prediction Research Manager, you will lead a focused team of Prediction Researchers and research engineers. You will translate Aaru's broader prediction agenda into a small number of important, testable workstreams and be accountable for the quality, pace, and practical impact of the team's research.
What they're looking for
- A new method improves average error but becomes badly miscalibrated during rapid change. Determine whether to model drift explicitly, add time-dependent conditioning, change the objective, or abstain when evidence is weak
- A language-model agent produces persuasive forecasts but does not outperform a historical-rate baseline. Design the ablations that reveal whether retrieval, decomposition, tool use, or reasoning contributes actual signal
- Aggregate predictions are accurate while subgroup estimates are unstable. Identify whether the failure comes from sparse data, biased sampling, weak population representations, excessive decomposition, or invalid evaluation
- A customer asks about a novel event with little direct historical precedent. Determine what adjacent evidence is transferable, how much uncertainty is irreducible, and whether a conditional scenario model can be validated honestly
- A method backtests well but fails prospectively. Investigate temporal leakage, benchmark selection, outcome revisions, implicit access to future information, and changes in the data-generating process
- Direct prediction and explicit population simulation disagree materially. Build a comparison that determines which representation, assumptions, and error sources explain the gap
More about this role
Aaru builds simulations of human behavior. Each simulation contains a population of AI agents, each representing a person who could plausibly exist in the real world and capable of making decisions within a modeled environment. Companies and institutions use these simulations to test consequential choices before committing—from product launches and pricing decisions to strategic communications and policy changes.
Building a useful simulation requires more than generating plausible text. Populations must represent real people and groups; predictions must be calibrated; simulations must remain coherent as conditions change; and the product must make the resulting evidence legible enough to support real decisions.
We are a small, in-person team in New York. We work with urgency, high ownership, and intellectual honesty. We expect people to surface inconvenient evidence, change their minds quickly, and carry important work all the way to a result.
Prediction Research builds systems that estimate future or otherwise unknown outcomes from data. The team's primary object is the population-level outcome: given a population, a question, and the relevant context, what aggregate result...
Browse similar: Startup jobs