Unstructured-IO/unstructured
Convert documents to structured data effortlessly. Unstructured is open-source ETL solution for transforming complex documents into clean, structured formats for language models. Visit our website to learn more about our enterprise grade Platform product for production grade workflows, partitioning, enrichments, chunking and embedding.
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 46%Set up a retrieval pipeline →
- PossiblePossibly related (embedding) · 45%Launch HN: Parsewise (YC P25) – Reason Across Documents with an API →
- PossiblePossibly related (embedding) · 45%The emergence of the web data infrastructure layer for AI →
- FuzzySimilar title/name (fuzzy) · 87%JobHop v2: A Large-Scale Career Trajectory Dataset from Unstructured Resumes →
“Fuzzy title match (0.94): “JobHop v2: A Large-Scale Career Trajectory Dataset from Unst” ≈ “Unstructured-IO/unstructured””
- FuzzySimilar title/name (fuzzy) · 87%Efficient Compression of Structured and Unstructured Volumes via Learned 3D Gaussian Representation →
“Fuzzy title match (0.94): “Efficient Compression of Structured and Unstructured Volumes” ≈ “Unstructured-IO/unstructured””
- FuzzySimilar title/name (fuzzy) · 87%DKCD: Domain Knowledge-Enhanced Causal Discovery from Unstructured Data →
“Fuzzy title match (0.94): “DKCD: Domain Knowledge-Enhanced Causal Discovery from Unstru” ≈ “Unstructured-IO/unstructured””
