The RAT: A Unified Bayesian Model for RAG Evaluation
Evaluating Retrieval-Augmented Generation (RAG) systems requires assessing not only end-to-end correctness but also how individual components interact and how errors propagate through the pipeline. We introduce a Bayesian evaluation framework that jointly models retrieval success, abstention behavior, and answer correctness, factorized according to the pipeline's information flow. The model distinguishes task success. Whether the user received a correct answer (from generator success) and whether the generator behaved appropriately given the retrieval outcome. We apply the framework to 27 RAG
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 50%Cutting RAG inference costs 6x starts with deciding what never reaches the LLM →
- PossiblePossibly related (embedding) · 26%NirDiamant/RAG_Techniques →
“Possibly related via embedding similarity 0.57 (not asserted). Timestamp check: artifact slightly before paper (-53d).”
- LinkedLinked via arxiv author · 85%Pius von Däniken →
“The RAT: A Unified Bayesian Model for RAG Evaluation”
- LinkedLinked via arxiv author · 85%Felix Matthias Saaro →
“The RAT: A Unified Bayesian Model for RAG Evaluation”
- LinkedLinked via arxiv author · 85%Mark Cieliebak →
“The RAT: A Unified Bayesian Model for RAG Evaluation”
- LinkedLinked via arxiv author · 85%Jan Deriu →
“The RAT: A Unified Bayesian Model for RAG Evaluation”
