Hindcast: Replaying Prediction Markets to Evaluate LLM Forecasters
Forecasters are evaluated by backtesting, which replays resolved questions and grades the probability the system would have assigned before the outcome was known. For LLMs, two channels leak the answer into this test. A model that retrieves can surface reports written after the event, turning forecasting into a lookup, and each new model is trained on data closer to the event, so a question that lay in the future for last year's models sits inside this year's training data. Either way, the test grades recall while claiming to grade foresight. We introduce Hindcast, which closes both leaks by g
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 46%Evaluate a model properly →
- FuzzyOverlapping authors or contributors · 62%sgl-project/sglang →
“Shared author/contributor keys: zhou”
- LinkedLinked via arxiv author · 85%Xiao Ye →
“Hindcast: Replaying Prediction Markets to Evaluate LLM Forecasters”
- LinkedLinked via arxiv author · 85%Jacob Dineen →
“Hindcast: Replaying Prediction Markets to Evaluate LLM Forecasters”
- LinkedLinked via arxiv author · 85%Evan Zhu →
“Hindcast: Replaying Prediction Markets to Evaluate LLM Forecasters”
- LinkedLinked via arxiv author · 85%Shijie Lu →
“Hindcast: Replaying Prediction Markets to Evaluate LLM Forecasters”
- LinkedLinked via arxiv author · 85%Kevin Song →
“Hindcast: Replaying Prediction Markets to Evaluate LLM Forecasters”
- LinkedLinked via arxiv author · 85%Ben Zhou →
“Hindcast: Replaying Prediction Markets to Evaluate LLM Forecasters”
