Read original ↗
paperarXivTrust 82 · PrimaryPublished 2mo agoLive · 1mo ago

Measuring the Gap Between Human and LLM Research Ideas

LLMs are increasingly used to brainstorm research ideas, but existing evaluations mostly judge individual ideas by novelty, feasibility, or expert preference. We instead ask: how far are current LLM-generated ideas from human researchers? To characterize this gap, we build a large-scale evaluation framework for ideation from high-quality human research papers. For each paper, we reverse-engineer a small set of closely related prior works that likely inspired its core idea. LLMs are then prompted to generate a new idea from the set of paper titles and summaries. We introduce a two-axis research

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • LinkedLinked via arxiv author · 85%Ziyu Chen

    Measuring the Gap Between Human and LLM Research Ideas

  • LinkedLinked via arxiv author · 85%Yilun Zhao

    Measuring the Gap Between Human and LLM Research Ideas

  • LinkedLinked via arxiv author · 85%Arman Cohan

    Measuring the Gap Between Human and LLM Research Ideas

  • PossiblePossibly related (embedding) · 45%lechmazur/writing

authored (incoming)

Implements (incoming)

Related across the graph

Topics