Think-Probe-Respond: Improving Large Language Models as Judges of Research Idea Novelty
Automated novelty judgment can accelerate scientific discovery by enabling efficient evaluation, refinement, and comparison of research ideas. While large language models are increasingly adopted for this task, we investigate a previously overlooked limitation in their judgment capabilities: despite generating reasoning rationales that closely mirror those of human experts, their final novelty judgments often diverge substantially. We demonstrate that this miscalibration stems from a systematic bias towards judging ideas as "medium novel". To mitigate this, we propose Think-Probe-Respond (TPR)
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 53%When AI agrees with us: large language models as amplifiers of scientific bias - Springer Nature Link →
- PossiblePossibly related (embedding) · 48%Understanding large language models demands distinguishing human projection from machine cognition - Nature →
- PossiblePossibly related (embedding) · 48%Researchers Explore What It Means To Say AI 'Thinks' - Digital Information World →
- PossiblePossibly related (embedding) · 47%When machines misread science: creating guardrails for human and AI interpretation of biomedical research - Nature →
- FuzzySimilar title/name (fuzzy) · 59%google-research/google-research →
“Fuzzy title match (0.73): “Think-Probe-Respond: Improving Large Language Models as Judg” ≈ “google-research/google-research””
- LinkedLinked via arxiv author · 85%Tim Schopf →
“Think-Probe-Respond: Improving Large Language Models as Judges of Research Idea Novelty”
- LinkedLinked via arxiv author · 85%Tobias Schreieder →
“Think-Probe-Respond: Improving Large Language Models as Judges of Research Idea Novelty”
- LinkedLinked via arxiv author · 85%Akiko Aizawa →
“Think-Probe-Respond: Improving Large Language Models as Judges of Research Idea Novelty”
