repoGitHubTrust 82 · PrimaryPublished 3d agoLive · 3d ago
Sriram-PR/doc-scraper
Go web crawler to scrape documentation sites and convert content to clean Markdown for LLM ingestion (RAG, training data).
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 52%LlamaIndex ‘legal-kb’: Agentic Retrieval over Index v2 with retrieve, find, read, and grep Tools - MarkTechPost →
- PossiblePossibly related (embedding) · 51%What's in your RAG? →
- PossiblePossibly related (embedding) · 51%Detecting LLM-Generated Texts with "Classical" Machine Learning →
- PossiblePossibly related (embedding) · 49%Pydantic + OpenAI: The Cleanest Way to Get Structured Outputs from LLMs - Towards Data Science →
- PossiblePossibly related (embedding) · 46%LLM Provenance: Tracking Data Origins with Graffiti - StartupHub.ai →
Covers
newsLlamaIndex ‘legal-kb’: Agentic Retrieval over Index v2 with retrieve, find, read, and grep Tools - MarkTechPostnewsWhat's in your RAG?newsDetecting LLM-Generated Texts with "Classical" Machine LearningnewsPydantic + OpenAI: The Cleanest Way to Get Structured Outputs from LLMs - Towards Data SciencenewsLLM Provenance: Tracking Data Origins with Graffiti - StartupHub.ai
Related across the graph
newsDetecting LLM-Generated Texts with "Classical" Machine LearningnewsPydantic + OpenAI: The Cleanest Way to Get Structured Outputs from LLMs - Towards Data SciencenewsLlamaIndex ‘legal-kb’: Agentic Retrieval over Index v2 with retrieve, find, read, and grep Tools - MarkTechPostnewsLLM Provenance: Tracking Data Origins with Graffiti - StartupHub.ainewsWhat's in your RAG?
