Prediction Research Manager
About Aaru
Aaru builds simulations of human behavior. Each simulation contains a population of AI agents, each representing a person who could plausibly exist in the real world and capable of making decisions within a modeled environment. Companies and institutions use these simulations to test consequential choices before committing—from product launches and pricing decisions to strategic communications and policy changes.
Building a useful simulation requires more than generating plausible text. Populations must represent real people and groups; predictions must be calibrated; simulations must remain coherent as conditions change; and the product must make the resulting evidence legible enough to support real decisions.
We are a small, in-person team in New York. We work with urgency, high ownership, and intellectual honesty. We expect people to surface inconvenient evidence, change their minds quickly, and carry important work all the way to a result.
About Prediction Research
Prediction Research builds systems that estimate future or otherwise unknown outcomes from data. The team's primary object is the population-level outcome: given a population, a question, and the relevant context, what aggregate result should we expect, how uncertain should we be, and how should that estimate change when the conditions change?
Some problems are best solved with structured statistical or machine-learning methods. Others may benefit from language models, retrieval, tools, explicit decomposition, simulated agents, or a combination of these approaches. The team's job is not to assume that the most complex method is best. It is to determine which information and method produce genuine predictive signal beyond strong, simpler baselines.
Prediction Research is not prompt engineering and it is not a speculative forecasting exercise. It is empirical predictive science. A prediction of 60 percent should resolve near 60 percent under the conditions where it is made. Improvements must survive temporal holdouts, new populations, changing environments, and prospective outcomes.
The role
As Prediction Research Manager, you will lead a focused team of Prediction Researchers and research engineers. You will translate Aaru's broader prediction agenda into a small number of important, testable workstreams and be accountable for the quality, pace, and practical impact of the team's research.
Managers at Aaru remain researchers. You will write code, design experiments, review statistical assumptions, inspect individual failures, and directly contribute to the hardest technical questions. You will also hire exceptional people, develop researchers, provide candid feedback, create clear ownership, and build an operating cadence that supports both fast exploration and rigorous final evidence.
You will work closely with Population Research, Evaluation Research, Simulation Engineering, Product Engineering, Research Product, Data, and Deployment. Population Research supplies representations of people and groups; Evaluation Research supplies protected measurements and diagnostic evidence; Simulation Engineering turns validated methods into dependable production capabilities. Your team must make these interfaces explicit and productive.
What you will do
Build, lead, and develop a high-performing team of Prediction Researchers and research engineers.
Set a focused portfolio of research questions across aggregate prediction, behavioral forecasting, calibration, conditioning, subgroup decomposition, drift, and agentic prediction systems.
Turn poorly specified prediction problems into falsifiable hypotheses, strong baselines, clear datasets, decisive experiments, and explicit criteria for continuing or stopping a line of work.
Build and test methods that combine structured data, statistical and machine-learning models, language models, retrieval, tools, and inference-time reasoning where each component earns its complexity.
Develop predictive systems from real-world records such as transactions, product usage, event histories, operational data, market data, surveys, customer data, and longitudinal outcomes.
Improve estimates of population behavior and how those estimates vary with attributes, prior behavior, information exposure, environment, time, and intervention.
Establish standards for calibration, proper scoring, temporal validity, prospective testing, subgroup performance, selective prediction, leakage prevention, and performance under condition shift.
Require credible comparisons with historical rates, conventional statistical models, direct prediction, segment-level methods, and population-based simulation.
Determine when modeling individual agents produces meaningful predictive value over direct aggregate estimation—and when it does not.
Build feedback loops in which resolved events and customer outcomes improve future systems while protected evaluation sets remain clean.
Partner with data and infrastructure teams on the acquisition, joining, cleaning, documentation, provenance, and reliable use of research-quality data.
Work with Population Research to distinguish errors caused by an inaccurate population from errors caused by the prediction method or its conditioning information.
Work with Evaluation Research to build diagnostic and final evaluations that reveal not only whether a method improved, but where, why, and under which conditions.
Transfer validated methods to Simulation Engineering with clear contracts, reproducible implementations, known limitations, and evidence that supports productionization.
Communicate negative results, unstable improvements, and weak evidence plainly. Stop attractive projects when the data no longer justifies them.
Recruit exceptional researchers, set clear expectations, give direct feedback, develop independent judgment, and address performance or ownership problems early.
Representative research and leadership problems
You might be responsible for situations such as:
A new method improves average error but becomes badly miscalibrated during rapid change. Determine whether to model drift explicitly, add time-dependent conditioning, change the objective, or abstain when evidence is weak.
A language-model agent produces persuasive forecasts but does not outperform a historical-rate baseline. Design the ablations that reveal whether retrieval, decomposition, tool use, or reasoning contributes actual signal.
Aggregate predictions are accurate while subgroup estimates are unstable. Identify whether the failure comes from sparse data, biased sampling, weak population representations, excessive decomposition, or invalid evaluation.
A customer asks about a novel event with little direct historical precedent. Determine what adjacent evidence is transferable, how much uncertainty is irreducible, and whether a conditional scenario model can be validated honestly.
A method backtests well but fails prospectively. Investigate temporal leakage, benchmark selection, outcome revisions, implicit access to future information, and changes in the data-generating process.
Direct prediction and explicit population simulation disagree materially. Build a comparison that determines which representation, assumptions, and error sources explain the gap.
Multiple prediction classes require different methods. Define the taxonomy and routing logic without creating a collection of opaque, unmaintainable special cases.
The team has many promising research directions but no shared learning loop. Establish a portfolio, ownership model, experiment review, and stopping discipline that concentrates effort on the highest-value uncertainties.
A research result is statistically positive but too small or fragile to matter in production. Make the decision legible and redirect the team without overstating the result.
How we work
We treat prediction as an empirical science. Progress is measured against future or otherwise held-out outcomes, with particular attention to calibration, behavioral shift, subgroup performance, data leakage, and the cost of different errors. Strong baselines matter, including simple historical rates and conventional statistical models.
Research is exploratory, but it must eventually change what Aaru can build or what the company believes. Papers, benchmarks, and prototypes can be valuable along the way. The central goal is to produce predictive systems that remain useful when they encounter new data, new customers, and real consequences.
Management in this function combines scientific judgment with direct people leadership. The manager should create clarity about which uncertainties matter, provide researchers enough context to own them independently, and preserve the discipline to publish nulls, revise methods, or stop work when the evidence demands it.
You might thrive in this role if
You have led predictive, forecasting, machine-learning, quantitative, or applied research in an environment with a high empirical bar.
You have developed or directed an original research agenda rather than only executing a predefined roadmap.
You have built predictive systems from messy, heterogeneous data and tested them against observed, temporally valid outcomes.
You are comfortable moving between research strategy, statistical reasoning, model design, data design, implementation, and detailed error analysis.
You understand the strengths and failure modes of language models and can combine them productively with structured data and quantitative methods.
You can turn a vague, consequential question into a sequence of experiments that resolves the most important uncertainties first.
You care deeply about calibration, temporal validity, selection effects, leakage, subgroup behavior, and robustness under condition shift.
You can prioritize a research portfolio, set stopping rules, and direct resources toward results likely to change the product or the company's beliefs.
You have managed or technically led strong researchers, give clear feedback, and develop independent scientific judgment in others.
You communicate results clearly and are willing to change direction when the evidence contradicts an attractive idea.
You want to work in person in New York with a team that moves quickly and treats empirical truth as the standard.
Strong candidates may also have
Work in time-series modeling, econometrics, decision science, quantitative social science, recommender systems, risk modeling, causal inference, market prediction, or probabilistic programming.
Experience with LLM agents, retrieval, tool use, post-training, synthetic environments, model-based reasoning, or inference-time scaling.
Experience building proprietary datasets, data products, acquisition programs, or learning systems that improve as outcomes resolve.
A record of research that improved an operational prediction system or changed a consequential product, policy, investment, or business decision.
Experience with prospective forecasting, prediction markets, demand modeling, experimentation, or decision-making under uncertainty.
Experience recruiting and mentoring unusually strong researchers, research engineers, data scientists, or quantitative practitioners.
Experience translating research into a production system through close partnership with engineering teams.
Candidates need not have
Prior experience in human-behavior simulation or a career spent exclusively in language-model research.
A PhD, provided you have equivalent evidence of rigorous predictive work and research leadership.
Managed managers or a large organization; this role is about leading a focused team and remaining directly involved in the research.
What success looks like
The team has a clear prediction research portfolio organized around a small number of important, testable questions.
New methods outperform strong baselines on clean historical and prospective outcomes, with well-understood limits and uncertainty.
Forecasts and other predictive outputs are calibrated, useful in real decisions, and able to abstain or widen uncertainty when the evidence is weak.
The team can explain which methods work for which prediction classes, populations, and conditions rather than relying on a single undifferentiated approach.
Data has clear provenance, protected evaluation boundaries, and direct compounding value for model development.
Population, Prediction, and Evaluation teams can distinguish and localize sources of simulation error rather than debating aggregate results without diagnosis.
Validated methods move into production efficiently and improve the quality of Aaru's simulations and customer-facing products.
Researchers grow into independent owners of important technical directions, and the team maintains an unusually high bar for rigor, speed, and honest communication.
Location and benefits
This role is based in New York City. Aaru is an in-person company, working five days a week in the office. Candidates should be located in the New York metropolitan area or open to relocation.
Aaru offers a competitive base salary, equity participation, comprehensive medical, vision, and dental coverage, visa sponsorship and relocation support, and other benefits and perks. Final compensation depends on level and experience and is set within Aaru's internal bands.