repoGitHubTrust 82 · PrimaryPublished 1mo agoLive · 1mo ago
run-llama/ParseBench
ParseBench - A Document Parsing Benchmark for AI Agents
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 61%ScarfBench: Benchmarking AI Agents for Enterprise Java Framework Migration →
- PossiblePossibly related (embedding) · 60%What AI benchmarks are not telling you →
- PossiblePossibly related (embedding) · 58%Introducing GeneBench-Pro →
- PossiblePossibly related (embedding) · 55%A new era for AI Search →
- PossiblePossibly related (embedding) · 54%Launch HN: Parsewise (YC P25) – Reason Across Documents with an API →
- PossiblePossibly related (embedding) · 58%PACE: A Proxy for Agentic Capability Evaluation →
- PossiblePossibly related (embedding) · 56%A$^{2}$utoLPBench: An Auto-Generated, Agent-Friendly LP Benchmark via Inverse-KKT Construction →
- PossiblePossibly related (embedding) · 55%ToolFailBench: Diagnosing Tool-Use Failures in LLM Agents →
Covers
Implements (incoming)
paperPACE: A Proxy for Agentic Capability EvaluationpaperA$^{2}$utoLPBench: An Auto-Generated, Agent-Friendly LP Benchmark via Inverse-KKT ConstructionpaperToolFailBench: Diagnosing Tool-Use Failures in LLM AgentspaperPolyWorkBench: Benchmarking Multilingual Long-Horizon LLM AgentspaperRuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task SpecificationspaperAn Experimental Design Approach to Evaluating Agentic AI's Autonomous Model Discovery
Covers (incoming)
newsGoogle Cloud tests AI agents with ambiguity-based benchmarks - IT Brief AustralianewsShow HN: Benchmark your eng team's AI agent maturity in 5 minutesnewsGerman AI consortium releases Soofi S, an open 30B model that tops benchmarksnewsPerplexity AI Releases WANDR: An Open Benchmark Evaluating Research Agents That Must Search Wide And Deep - MarkTechPost
Related across the graph
newsPerplexity AI Releases WANDR: An Open Benchmark Evaluating Research Agents That Must Search Wide And Deep - MarkTechPostpaperPolyWorkBench: Benchmarking Multilingual Long-Horizon LLM AgentspaperAn Experimental Design Approach to Evaluating Agentic AI's Autonomous Model DiscoverynewsShow HN: Benchmark your eng team's AI agent maturity in 5 minutesnewsWhat AI benchmarks are not telling younewsLaunch HN: Parsewise (YC P25) – Reason Across Documents with an APIpaperPACE: A Proxy for Agentic Capability EvaluationpaperToolFailBench: Diagnosing Tool-Use Failures in LLM AgentsnewsScarfBench: Benchmarking AI Agents for Enterprise Java Framework MigrationnewsA new era for AI SearchpaperRuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task SpecificationsnewsGoogle Cloud tests AI agents with ambiguity-based benchmarks - IT Brief AustraliapaperA$^{2}$utoLPBench: An Auto-Generated, Agent-Friendly LP Benchmark via Inverse-KKT ConstructionnewsGerman AI consortium releases Soofi S, an open 30B model that tops benchmarksnewsIntroducing GeneBench-Pro
