repoGitLabTrust 82 · PrimaryPublished 3d agoLive · 3d ago
santushtmatra11/structured-document-retriever
A modular Retrieval-Augmented Generation (RAG) system for textbooks and long-form documents. Uses layout-aware heuristics + LLM-verified chunking to generalize across document formats (no hardcoded regex), stores embeddings and content in PostgreSQL/pgvector, and answers questions with query-routing guardrails (greeting/in-domain/off-topic/blocked detection).
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 61%Set up a retrieval pipeline →
- PossiblePossibly related (embedding) · 56%A Production RAG Pipeline for PDFs: Relational Parsing, TOC Retrieval, Typed Answers - Towards Data Science →
