Read original ↗
paperarXivTrust 82 · PrimaryPublished 1mo agoLive · 1mo ago

Hindcast: Replaying Prediction Markets to Evaluate LLM Forecasters

Forecasters are evaluated by backtesting, which replays resolved questions and grades the probability the system would have assigned before the outcome was known. For LLMs, two channels leak the answer into this test. A model that retrieves can surface reports written after the event, turning forecasting into a lookup, and each new model is trained on data closer to the event, so a question that lay in the future for last year's models sits inside this year's training data. Either way, the test grades recall while claiming to grade foresight. We introduce Hindcast, which closes both leaks by g

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • PossiblePossibly related (embedding) · 46%Evaluate a model properly
  • FuzzyOverlapping authors or contributors · 62%sgl-project/sglang

    Shared author/contributor keys: zhou

  • LinkedLinked via arxiv author · 85%Xiao Ye

    Hindcast: Replaying Prediction Markets to Evaluate LLM Forecasters

  • LinkedLinked via arxiv author · 85%Jacob Dineen

    Hindcast: Replaying Prediction Markets to Evaluate LLM Forecasters

  • LinkedLinked via arxiv author · 85%Evan Zhu

    Hindcast: Replaying Prediction Markets to Evaluate LLM Forecasters

  • LinkedLinked via arxiv author · 85%Shijie Lu

    Hindcast: Replaying Prediction Markets to Evaluate LLM Forecasters

  • LinkedLinked via arxiv author · 85%Kevin Song

    Hindcast: Replaying Prediction Markets to Evaluate LLM Forecasters

  • LinkedLinked via arxiv author · 85%Ben Zhou

    Hindcast: Replaying Prediction Markets to Evaluate LLM Forecasters

Explains

Implements (incoming)

authored (incoming)

Related across the graph

Topics