forecasting · activePredicting future events
Whether a model can make well-calibrated predictions about events that haven't happened yet, not just recall or verify existing facts.
Tags: reasoning, knowledge
Distinct from fact-verification, which checks claims against existing evidence. Forecasting requires synthesizing current information toward a genuinely uncertain future outcome, and calibration — knowing how confident to be, not just landing on the right answer — is the real test. Genuinely hard to contaminate, since the events haven't happened at training time, which is also what makes it hard to know whether poor performance reflects the event being unpredictable or the model being bad at synthesizing evidence.
Claims
- observationsingle paperA retrieval-augmented forecasting system built on a GPT-4-class model — not the base model alone — approaches, and sometimes exceeds, the accuracy of competitive human forecasters on real prediction-market questions dated after the model's training cutoff.
- observationsingle paperWithout added retrieval infrastructure, language models underperform human experts at forecasting real-world events from a benchmark of actual forecasting-tournament questions, though accuracy improves with model scale and access to relevant news context.
Techniques
None yet.
Capabilities are a way of carving up the subject, and carvings are arguable. Say so if this one is wrong — especially a proposed one, which a pipeline added because several papers used the same framing, not because anyone decided it was right.