Before the Action: Benchmarking LLMs on Prospective Hypothesis Discovery
Large language models (LLMs) excel at answering pre-specified questions, yet their ability to navigate the open-ended, pre-conclusion stage of discovery remains largely unmeasured. We introduce Prospective Hypothesis Discovery (PHD), which asks models to autonomously construct grounded, discriminative, and testable hypothesis spaces from inconclusive evidence, including anomalous observations and fragmented records, to guide subsequent investigation. To evaluate this capability, we introduce HypoArena, comprising HypoData, a benchmark of 988 cases across six scientific and analytical domains,
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 53%Understanding large language models demands distinguishing human projection from machine cognition - Nature →
- PossiblePossibly related (embedding) · 52%Large Language Models Are Still Getting Stronger, but Researchers Face New Bottlenecks in Data, Evaluation, and Safety | Newswise - Newswise →
- FuzzySimilar title/name (fuzzy) · 84%liguodongiot/llm-action →
“Fuzzy title match (0.92): “Before the Action: Benchmarking LLMs on Prospective Hypothes” ≈ “liguodongiot/llm-action””
- FuzzyOverlapping authors or contributors · 62%affaan-m/ECC →
“Shared author/contributor keys: jiang”
- FuzzyOverlapping authors or contributors · 62%google-research/google-research →
“Shared author/contributor keys: sun”
- FuzzyOverlapping authors or contributors · 62%ultralytics/ultralytics →
“Shared author/contributor keys: han”
- FuzzyOverlapping authors or contributors · 62%Zeyi-Lin/HivisionIDPhotos →
“Shared author/contributor keys: lin”
- LinkedLinked via arxiv author · 85%Tianyun Zhong →
“Before the Action: Benchmarking LLMs on Prospective Hypothesis Discovery”
