Bridge Evidence: Static Retrieval Utility Does Not Predict Causal Utility in Multi-Step Agentic Search
Retrieval systems are trained and evaluated on a static idea of usefulness: hand a document and a question to a reader model, see whether the answer improves, and score the document accordingly. The idea holds up when a document is read on its own. It breaks when a language model works as a search agent, issuing several queries and reasoning across turns, because a document can matter for what it lets the agent do next rather than for what it says about the current question. We measure that gap rather than argue it. Using a ReAct style agent over HotpotQA, we replay 1000 development question
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 53%CLaRa: Bridging Retrieval and Generation with Continuous Latent Reasoning - Apple Machine Learning Research →
- PossiblePossibly related (embedding) · 52%Agentic Resource Discovery: Let agents search →
- FuzzySimilar title/name (fuzzy) · 59%Fosowl/agenticSeek →
“Fuzzy title match (0.73): “Bridge Evidence: Static Retrieval Utility Does Not Predict C” ≈ “Fosowl/agenticSeek””
- LinkedLinked via arxiv author · 85%Debayan Mukhopadhyay →
“Bridge Evidence: Static Retrieval Utility Does Not Predict Causal Utility in Multi-Step Agentic Search”
- LinkedLinked via arxiv author · 85%Utshab Kumar Ghosh →
“Bridge Evidence: Static Retrieval Utility Does Not Predict Causal Utility in Multi-Step Agentic Search”
- LinkedLinked via arxiv author · 85%Shubham Chatterjee →
“Bridge Evidence: Static Retrieval Utility Does Not Predict Causal Utility in Multi-Step Agentic Search”
