PathAgentBench: Benchmarking Evidence-Seeking Vision-Language Models on Whole-Slide Pathology Image
Whole-slide image (WSI) diagnosis requires identifying diagnostically relevant regions, examining them across magnifications, and integrating multi-scale evidence. However, most existing pathology benchmarks evaluate models on pre-cropped patches or pre-extracted slide features, leaving their ability to acquire evidence directly from gigapixel WSIs largely untested. We introduce PathAgentBench, a benchmark for evaluating evidence-seeking vision-language models (VLMs) across four complementary capabilities: image-to-text matching for evidence interpretation, text-to-image retrieval for evidence
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- FuzzySimilar title/name (fuzzy) · 59%AgentCore-8B →
“Fuzzy title match (0.73): “PathAgentBench: Benchmarking Evidence-Seeking Vision-Languag” ≈ “AgentCore-8B””
- FuzzySimilar title/name (fuzzy) · 59%Tongyi-MAI/Z-Image-Turbo →
“Fuzzy title match (0.73): “PathAgentBench: Benchmarking Evidence-Seeking Vision-Languag” ≈ “Tongyi-MAI/Z-Image-Turbo””
- FuzzySimilar title/name (fuzzy) · 59%VioletVision-3B →
“Fuzzy title match (0.73): “PathAgentBench: Benchmarking Evidence-Seeking Vision-Languag” ≈ “VioletVision-3B””
- PossiblePossibly related (embedding) · 48%Benchmarking large language models against practicing clinicians on psychopathological assessment - Nature →
- FuzzySimilar title/name (fuzzy) · 87%SWE-agent/SWE-agent →
“Fuzzy title match (0.94): “PathAgentBench: Benchmarking Evidence-Seeking Vision-Languag” ≈ “SWE-agent/SWE-agent””
- FuzzySimilar title/name (fuzzy) · 87%zhayujie/CowAgent →
“Fuzzy title match (0.94): “PathAgentBench: Benchmarking Evidence-Seeking Vision-Languag” ≈ “zhayujie/CowAgent””
- FuzzySimilar title/name (fuzzy) · 84%pytorch/vision →
“Fuzzy title match (0.92): “PathAgentBench: Benchmarking Evidence-Seeking Vision-Languag” ≈ “pytorch/vision””
- FuzzyOverlapping authors or contributors · 62%keras-team/keras →
“Shared author/contributor keys: jin”
