articlecommunityTrust 52 · CommunityPublished 1mo agoLive · 3mo ago
Retrieval is underrated
Most 'reasoning' failures are really retrieval failures in disguise.
Most 'reasoning' failures are really retrieval failures in disguise. Most 'reasoning' failures are really retrieval failures in disguise.
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- LinkedLinked via unknownSelf-rewarding agents that retrace failures →
- LinkedLinked via unknownCORTEX: A Structured Reasoning Benchmark for Trustworthy 3D Chest CT MLLMs →
- LinkedLinked via unknownThe Riddle Riddle: Testing Flexible Reasoning in Large Language Models and Humans →
- LinkedLinked via unknownRetrace-1.5B →
- LinkedLinked via unknownNew benchmark exposes reasoning gaps in top models →
- LinkedLinked via unknownEvidence-Informed LLM Beliefs for Continual Scientific Discovery →
- LinkedLinked via unknownModality-Driven Search with Holistic Trace Judging for ARC-AGI-2 →
Related to
Related to (incoming)
paperCORTEX: A Structured Reasoning Benchmark for Trustworthy 3D Chest CT MLLMspaperThe Riddle Riddle: Testing Flexible Reasoning in Large Language Models and HumansmodelRetrace-1.5BpaperCognitive Episodes in LLM Reasoning Traces Enable Interpretable Human Item Difficulty PredictionpaperEvidence-Informed LLM Beliefs for Continual Scientific DiscoverypaperModality-Driven Search with Holistic Trace Judging for ARC-AGI-2paperCheckRLM: Effective Knowledge-Thought Coherence Checking in Retrieval-Augmented ReasoningpaperDynaKRAG: A Unified Framework for Learnable Evidence Control in Multi-Hop Retrieval-Augmented GenerationpaperEvidence Interfaces Shape How Retrieval-Augmented Readers Use Support
Covers (incoming)
Related across the graph
paperEvidence Interfaces Shape How Retrieval-Augmented Readers Use SupportpaperCheckRLM: Effective Knowledge-Thought Coherence Checking in Retrieval-Augmented ReasoningpaperDynaKRAG: A Unified Framework for Learnable Evidence Control in Multi-Hop Retrieval-Augmented GenerationpaperSelf-rewarding agents that retrace failurespaperEvidence-Informed LLM Beliefs for Continual Scientific DiscoverymodelRetrace-1.5BpaperCORTEX: A Structured Reasoning Benchmark for Trustworthy 3D Chest CT MLLMspaperModality-Driven Search with Holistic Trace Judging for ARC-AGI-2paperThe Riddle Riddle: Testing Flexible Reasoning in Large Language Models and HumansnewsNew benchmark exposes reasoning gaps in top modelspaperCognitive Episodes in LLM Reasoning Traces Enable Interpretable Human Item Difficulty Prediction
