repoGitHubTrust 82 · PrimaryPublished 27d agoLive · 27d ago
agentscope-ai/PawBench
A benchmark for evaluating LLM × harness performance.
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 54%Open Source Local LLM Training Tool (for consumer hardware) →
- PossiblePossibly related (embedding) · 51%New LLM Coordination Benchmark - Benchmarking Open-Ended Multi-Agent Coordination in Language Agents [R] →
- PossiblePossibly related (embedding) · 50%[Research Article] An LLM-based multi-agent system for remote sensing analysis - EurekAlert! →
- PossiblePossibly related (embedding) · 54%Training a harness for model-agnostic and task-environment-agnostic capability improvements with PyTorch-like framework [P] →
Covers
Covers (incoming)
Related across the graph
newsOpen Source Local LLM Training Tool (for consumer hardware)newsTraining a harness for model-agnostic and task-environment-agnostic capability improvements with PyTorch-like framework [P]newsNew LLM Coordination Benchmark - Benchmarking Open-Ended Multi-Agent Coordination in Language Agents [R]news[Research Article] An LLM-based multi-agent system for remote sensing analysis - EurekAlert!
